OpenAI GPT-6.1 Sol System CardProvider report
2026-09-29Twelve capability metrics from the GPT-6.1 Sol system card, with same-source GPT-6 Sol and GPT-6 Astra comparisons. Four HealthBench metrics use length-adjusted scores, retaining unadjusted scores and response lengths separately. Four MentalHealthBench metrics retain maximum reasoning effort, standard errors, and task counts. Four cybersecurity evaluations preserve metric definitions and historical-vulnerability contamination caveats; reasoning effort is recorded only when explicit. Comparison models may reflect later versions; historical release snapshots are unchanged.