Claude Opus 5.5 Official Release EvaluationProvider report
2026-09-22The September 22, 2026 release table and System Card capability summary on page 174, plus DeepSWE on page 175 and Chartography without tools on page 200: 18 metrics and 66 scores after deduplication. Missing cells create no scores. Opus 5.5 uses adaptive thinking at max effort except Terminal-Bench 4.0 at xhigh; GPT-6 Astra uses high on that test. OSWorld 2.0 uses September 10 tasks and server-side context compaction, with partial and strict metrics kept separate from older versions. AutomationBench uses Zapier's private held-out set with safeguard interventions counted as failures and no fallback. Other Opus 5.5 evaluations retain the reported production safeguards and fallbacks. FrontierCode keeps 54.4 at max as the main value and 54.6 at medium as supplementary data. Comparison scores retain this vendor's provenance and do not replace historical reports.