Terminal-Bench v3.0Coding
Official public leaderboard; each model uses the highest-scoring thinking level reported by the Terminal Bench 3.0 authors. Gemini results are self-computed with a mini swe agent (LiteLLM 1.96).Terminal-Bench 3.0 is the third-generation terminal-agent benchmark in the Terminal-Bench series. Its tasks cover general agent capabilities and evaluate how models autonomously complete complex tasks in real terminal environments. Each model's score uses the highest-scoring thinking level reported by the benchmark authors.