GPT-6 Astra
As of 14 Sep 2026, GPT-6 Astra is the RightSignal pick for best LLM for coding.
GPT-6 Astra tops Artificial Analysis's independent Terminal-Bench v4.0 run at 59.6% on xhigh and 59.1% on max, with Claude Fable 5.1 next at 55.1% and Opus 5 absent from the published top three.
Still the leader on Vals.ai's independent SWE-bench Verified run at 97.00%, against the 96.0% Anthropic self-reports in the Opus 5 system card, on a board Vals has now archived as saturated and never ran GPT-6 Astra on. On Vals' Terminal-Bench 2.1 it sits fourth at 84.6%, behind Astra, GPT-5.6 Sol and Claude Fable 5.1, and it does not appear in the published Terminal-Bench 4.0 top three.
Why it’s the challengercurrent reign
held off
reigns
Why GPT-6 Astra is the pick
- GPT-6 Astra tops Artificial Analysis's independent Terminal-Bench v4.0 run at 59.6% on xhigh and 59.1% on max, with Claude Fable 5.1 next at 55.1% and Opus 5 absent from the published top three.
- Vals.ai has it first on Terminal-Bench 2.1 too, at 87.3% ahead of GPT-5.6 Sol (85.8%), Fable 5.1 (85.0%) and Opus 5 (84.6%).
- It has no independent SWE-bench Verified run — Vals archived that board as saturated — and it lists at $10/$50 per million tokens, 2.5x GPT-5.6 Sol, so price the high-volume repo-scale batches carefully.
Evidence
The sources behind this title’s record.
- benchlm.ai/benchmarks/swe-bench-verified
- developers.openai.com/api/docs/models/gpt-6-astra
- www.tbench.ai/leaderboard/terminal-bench/2.1
- arena.ai/leaderboard/code/webdev
- www.anthropic.com/claude/opus
- openai.com/index/codex-quantum-computing-experiments
- arxiv.org/abs/2609.02272
- www.anthropic.com/news/claude-opus-5
Vendor numbers are treated as claims until independently reproduced — how we judge. Structured benchmark comparisons are on the roadmap.
Caveats & challengers
- vals.ai's Terminal-Bench 2.1 table now has Opus 5 fourth at 84.64%, behind GPT-6 Astra (87.27%), GPT-5.6 Sol (85.77%) and Claude Fable 5.1 (85.02%), and its own figure drops to 81.27% if you count its server-side fallbacks to Opus 4.8 on refused tasks as failures.
- The SWE-bench Verified lead is thin as well: DeepSeek V4 Pro (0813) sits second at 96.40% against 97.00%, as an open-weight model costing roughly $0.02 per task against $1.29 for Opus 5.
- Worth knowing that vals has since dropped SWE-bench Verified from its coding index altogether, calling the benchmark saturated, so lean on terminal and repo-scale evidence rather than that one number.
Still the leader on Vals.ai's independent SWE-bench Verified run at 97.00%, against the 96.0% Anthropic self-reports in the Opus 5 system card, on a board Vals has now archived as saturated and never ran GPT-6 Astra on. On Vals' Terminal-Bench 2.1 it sits fourth at 84.6%, behind Astra, GPT-5.6 Sol and Claude Fable 5.1, and it does not appear in the published Terminal-Bench 4.0 top three.
At a glance
- Title
- Best LLM for coding
- Licence
- proprietary
- Pick since
- 9 Sep 2026
- Last reviewed
- 13 Sep 2026
- Title holders to date
- 10
- Official page
- developers.openai.com/api/docs/mode...
Title history
Every change, on the record.
GPT-6 Astra tops Artificial Analysis's independent Terminal-Bench v4.0 run at 59.6% on xhigh and 59.1% on max, with Claude Fable 5.1 next at 55.1% and Opus 5 absent from the published top three.
5 daysas pick
New state of the art on SWE-bench Verified and the WebDev arena at half Fable 5's price — which made it the practical default for long-ru…
47 daysas pick
On global redeploy it scored highest of any frontier model on Cognition's FrontierCode even at medium effort.
23 daysas pick
Pushed SWE-bench Verified to 81.42% and took the top Terminal-Bench 2.0 score; the 4.7 and 4.8 refreshes extended the same era rather tha…
146 daysas pick
Took the crown up a tier — leading SWE-bench Verified and winning seven of eight languages on SWE-bench Multilingual.
73 daysas pick
Reclaimed the title decisively at 77.2% SWE-bench Verified, and jumped OSWorld computer use from 42.2% to 61.4% in four months.
56 daysas pick
Billed as the world's best coding model on 72.5% SWE-bench Verified, and held the practitioner crown through GPT-5's August challenge — t…
130 daysas pick
The first hybrid reasoning model, hitting 70.3% on SWE-bench Verified — and it shipped Claude Code alongside, which moved the goalposts f…
87 daysas pick
Flipped practitioner consensus away from OpenAI overnight by solving 64% of agentic coding problems against Claude 3 Opus's 38%; the Octo…
249 daysas pick
as pick