OpenAI GPT-6 Sol and Luna ReleaseProvider report
2026-09-22Five explicit scores across four benchmarks from the GPT-6 Sol and Luna release article. AutomationBench uses the v1.0.6 private held-out set; OSWorld uses partial reward on the v2026.08.08 offline set. Results retain their stated max or xhigh effort; xhigh is not labelled as max. Charts without readable exact values are not estimated, missing cells create no scores. One Astra low result from the same table is included for comparison within this AutomationBench version, for six scores in total. GPT evaluations use OpenAI's research environment or API and may differ from ChatGPT.