Claude Opus 5
Tops vals.ai's independent SWE-bench Verified run at 97.0%, the best of 82 systems on that board, and sits second on their Terminal-Bench 2.1 table at 84.64%. At $5/$25 per million tokens it is half the price of Claude Fable 5 ($10/$50) for scores within a point or two across the coding suite. Tracker refreshes through mid-August still name it the sensible default for repo-scale agentic work.
Challenger: Claude Fable 5 — Still tops the official Terminal-Bench 2.1 board at 83.8% with Claude Code, ahead of Codex on GPT-5.5 at 83.1%. That run dates from 7 June and no top-five entry post-dates Opus 5's 24 July launch — Opus 5 has no submitted run on that board at all, so its absence is a submission gap rather than a weaker score. Treat the Anthropic terminal figures with care either way: Vals discloses that both Fable 5's and Opus 5's runs on its harness used Opus 4.8 as a refusal fallback.
Genuinely contested on terminal work: vals.ai puts GPT-5.6 Sol first on Terminal-Bench 2.1 at 85.77% against Opus 5's 84.64%, and Opus 5's figure falls to 81.27% if you count its server-side fallbacks to Opus 4.8 on refused tasks as failures. The SWE-bench Verified lead is narrowing too — DeepSeek V4 Pro (0813) now sits second at 96.40% against 97.0%, as an open-weight model at a fraction of the cost per task.