Claude Sonnet 5.5 Official Release EvaluationProvider report
2026-09-28Anthropic's September 28, 2026 release table, System Card summary on page 109, and DeepSWE, Terminal-Bench-Science, Chartography and OSWorld sections: 18 metrics and 55 scores; missing cells create no scores. Default adaptive/max, except Opus 5.5 Terminal-Bench 4.0 at xhigh. FrontierCode keeps 46.2 at max, with 52.1 at xhigh as supplementary data. OSWorld 2.1 reuses the September 10 tasks and configuration already catalogued; partial and strict remain separate. HealthBench Professional is length-adjusted. AutomationBench uses the v1.0.6 private set with default fallbacks for Sonnet 5.5 and fallback reruns of refused Opus 5.5 tasks. Both AA Sonnet 5.5 scores came from a pre-release deployment with a structured-output bug, since fixed; any effect is expected to be small and to understate performance. Published scores and vendor provenance are preserved without replacing historical reports.