OpenAI GPT-6 Astra ReleaseProvider report
2026-09-03Based on the complete Evaluation comparison table on OpenAI's September 3, 2026 GPT-6 Astra release page, covering 40 metrics and 149 available scores; blank cells do not create scores. OpenAI states that table values are the maximum score at any reasoning effort and that GPT models were evaluated in its research environment or through the API, so results may differ slightly from production ChatGPT. The five AutomationBench scores use the official leaderboard's private held-out evaluation set and are stored separately as AutomationBench (Private), rather than being combined with scores from the public 600-task set. The OpenAI release page does not state the evaluation version for that row, so its original values are retained as a release snapshot. OSWorld, BenchCAD, and SRE-Bench are stored separately using the task version, tool configuration, and pass@1 metric stated on the release page. Five of the six safety metrics marked lower is better retain that direction. The lower-cost settings, SRE-Bench pass@4, and OSWorld timing results reported in the article are preserved in supplemental raw fields.