Meta Muse Spark 1.2 Evaluation ReportProvider report
2026-08-05Based on the five official evaluation charts in Meta's release announcement: Terminal-Bench 2.1 and DeepSWE 1.1 are self-computed by Meta (full official 89 / 113 task sets; each model runs under its own selected agent product at its highest available reasoning effort, with Muse Code for Muse Spark 1.2; averaged pass@1 over five attempts); GDPVal-AA v2 uses Artificial Analysis results; MCP Atlas uses Scale AI results; Meta Internal Coding Bench is an internal Meta benchmark (440 tasks, not externally reproducible). Values are taken from the official chart labels. Methodology: https://research.meta.ai/static/muse-spark-1-2-methodology