OpenAI GPT-6.1 Sol ReleaseProvider report
2026-09-29Ten metrics from the original GPT-6.1 Sol release charts, with GPT-6 Sol and GPT-6 Astra comparisons from the same charts. Each metric shows the best published result across tested efforts (minimum for error rates), with the actual effort labeled. All effort scores and per-task costs are retained in raw data. AutomationBench uses the v1.0.6 private set; OSWorld uses partial reward on the v2026.08.08 offline set. The five error or safety failure rates are lower-is-better; true zeros are retained. The computer-use safety stress test uses an updated, harder subset, separate from earlier safety metrics. Evaluations use the research environment or API; difficult-prompt error rates do not represent typical usage.