right signal
Pick Best LLM for coding

GPT-6 Astra

As of 14 Sep 2026, GPT-6 Astra is the RightSignal pick for best LLM for coding.

GPT-6 Astra tops Artificial Analysis's independent Terminal-Bench v4.0 run at 59.6% on xhigh and 59.1% on max, with Claude Fable 5.1 next at 55.1% and Opus 5 absent from the published top three.

Current pick
since 9 Sep 2026
5
days as pick
Next best (challenger)
Claude Opus 5

Still the leader on Vals.ai's independent SWE-bench Verified run at 97.00%, against the 96.0% Anthropic self-reports in the Opus 5 system card, on a board Vals has now archived as saturated and never ran GPT-6 Astra on. On Vals' Terminal-Bench 2.1 it sits fourth at 84.6%, behind Astra, GPT-5.6 Sol and Claude Fable 5.1, and it does not appear in the published Terminal-Bench 4.0 top three.

Why it’s the challenger
Reign history
5days
current reign
2challenges
held off
9 SEP 2026became pick
0previous
reigns
Last reviewed: 13 Sep 2026 Reviews are continuous. This pick can change when the evidence changes.

Why GPT-6 Astra is the pick

  • GPT-6 Astra tops Artificial Analysis's independent Terminal-Bench v4.0 run at 59.6% on xhigh and 59.1% on max, with Claude Fable 5.1 next at 55.1% and Opus 5 absent from the published top three.
  • Vals.ai has it first on Terminal-Bench 2.1 too, at 87.3% ahead of GPT-5.6 Sol (85.8%), Fable 5.1 (85.0%) and Opus 5 (84.6%).
  • It has no independent SWE-bench Verified run — Vals archived that board as saturated — and it lists at $10/$50 per million tokens, 2.5x GPT-5.6 Sol, so price the high-volume repo-scale batches carefully.
Judged on agentic coding benchmarks · repo-scale task performance · tool-calling reliability

Evidence

The sources behind this title’s record.

Vendor numbers are treated as claims until independently reproduced — how we judge. Structured benchmark comparisons are on the roadmap.

Caveats & challengers

  • vals.ai's Terminal-Bench 2.1 table now has Opus 5 fourth at 84.64%, behind GPT-6 Astra (87.27%), GPT-5.6 Sol (85.77%) and Claude Fable 5.1 (85.02%), and its own figure drops to 81.27% if you count its server-side fallbacks to Opus 4.8 on refused tasks as failures.
  • The SWE-bench Verified lead is thin as well: DeepSeek V4 Pro (0813) sits second at 96.40% against 97.00%, as an open-weight model costing roughly $0.02 per task against $1.29 for Opus 5.
  • Worth knowing that vals has since dropped SWE-bench Verified from its coding index altogether, calling the benchmark saturated, so lean on terminal and repo-scale evidence rather than that one number.
Challenger Claude Opus 5

Still the leader on Vals.ai's independent SWE-bench Verified run at 97.00%, against the 96.0% Anthropic self-reports in the Opus 5 system card, on a board Vals has now archived as saturated and never ran GPT-6 Astra on. On Vals' Terminal-Bench 2.1 it sits fourth at 84.6%, behind Astra, GPT-5.6 Sol and Claude Fable 5.1, and it does not appear in the published Terminal-Bench 4.0 top three.

At a glance

Title
Best LLM for coding
Licence
proprietary
Pick since
9 Sep 2026
Last reviewed
13 Sep 2026
Title holders to date
10
Official page
developers.openai.com/api/docs/mode...

Title history

Every change, on the record.

View full changelog
Pick 9 Sep 2026 – presentGPT-6 Astra Current

GPT-6 Astra tops Artificial Analysis's independent Terminal-Bench v4.0 run at 59.6% on xhigh and 59.1% on max, with Claude Fable 5.1 next at 55.1% and Opus 5 absent from the published top three.

5 days
as pick
Pick 24 Jul 2026 – 9 Sep 2026Claude Opus 5

New state of the art on SWE-bench Verified and the WebDev arena at half Fable 5's price — which made it the practical default for long-ru…

47 days
as pick
Pick 1 Jul 2026 – 24 Jul 2026Claude Fable 5

On global redeploy it scored highest of any frontier model on Cognition's FrontierCode even at medium effort.

23 days
as pick
Pick 5 Feb 2026 – 1 Jul 2026Claude Opus 4.6

Pushed SWE-bench Verified to 81.42% and took the top Terminal-Bench 2.0 score; the 4.7 and 4.8 refreshes extended the same era rather tha…

146 days
as pick
Pick 24 Nov 2025 – 5 Feb 2026Claude Opus 4.5

Took the crown up a tier — leading SWE-bench Verified and winning seven of eight languages on SWE-bench Multilingual.

73 days
as pick
Pick 29 Sep 2025 – 24 Nov 2025Claude Sonnet 4.5

Reclaimed the title decisively at 77.2% SWE-bench Verified, and jumped OSWorld computer use from 42.2% to 61.4% in four months.

56 days
as pick
Pick 22 May 2025 – 29 Sep 2025Claude Opus 4

Billed as the world's best coding model on 72.5% SWE-bench Verified, and held the practitioner crown through GPT-5's August challenge — t…

130 days
as pick
Pick 24 Feb 2025 – 22 May 2025Claude 3.7 Sonnet

The first hybrid reasoning model, hitting 70.3% on SWE-bench Verified — and it shipped Claude Code alongside, which moved the goalposts f…

87 days
as pick
Pick 20 Jun 2024 – 24 Feb 2025Claude 3.5 Sonnet

Flipped practitioner consensus away from OpenAI overnight by solving 64% of agentic coding problems against Claude 3 Opus's 38%; the Octo…

249 days
as pick
Pick 1 Aug 2023 – 20 Jun 2024GPT-4 324 days
as pick