right signal
141 entries · dated, cited, on the record

Changelog

Every change to the board back to 2023 — who took each title, from whom, and on what evidence. Entries marked retrospective were researched as part of the historical backfill; everything else was recorded as it happened.

13 Sep 2026
challenged

Wan 3.0's lead over Gemini Omni Flash widens to five Elo on AA's with-audio text-to-video board

Artificial Analysis's with-audio text-to-video board, read today, puts Alibaba Wan 3.0 first on 1242 with Gemini Omni Flash second on 1237, then fal's post-trained Minimax H3 Max on 1231 and MiniMax H3 Open Weights on 1225 (AA text-to-video). That is the third successive read with Wan 3.0 at or ahead of the pick — level at 1239, then 1240 to 1239, now five points clear — so this is no longer a one-off reading, and the direction of travel is against the pick.

We are not moving the title yet, for two reasons. First, the lead is confined to one board from one provider: the pick still tops AA's no-audio text-to-video table, where MiniMax H3 is second on 1301, and leads on image-to-video without audio on 1365. Five Elo points on a single with-audio table is thin corroboration against that.

Second, two of the four things this slot judges on — character consistency and clip length — still have no independent read for Wan 3.0 at all. We would be swapping a model whose limits we know (Google's own card lists complete consistency across edits as open; 720p standard output with upscaling for higher resolutions) for one whose limits are simply unmeasured.

If Wan 3.0 holds this lead into the next read, or picks up a second board or an independent consistency or clip-length result, the title moves.

evidence: [1] [2]
13 Sep 2026
challenged

Breeze TTS 2 slips to #9 as StepAudio 2.5 TTS edges ahead

The pick has lost ground. On the Artificial Analysis provider-voice board (fetched 13 September) BreezeBlue Breeze TTS 2 Open Weights reads 1202 Elo at #9, down from the 1215 and 7th place that won it the title on 9 September. Sitting just above it at 1207 is StepFun StepAudio 2.5 TTS (Aug 2026), with the rest of the top ten — Cartesia Sonic 3.6 (1276), Inworld Realtime TTS-2 (1243), Speechify Simba 3.2 (1237), Qwen-Audio-3.0-TTS-Plus (1234), VUI Luna (1228), Gemini 3.1 Flash TTS, ElevenLabs v3 — either hosted-only or unconfirmed for downloadable weights.

StepFun is the obvious challenger: the Step-Audio line already appears in this slot's caveats as the Apache-2.0-coded option self-hosters fall back on. But we have nothing in front of us confirming that checkpoints for the 2.5 TTS release have actually shipped, nor its licence terms, time-to-first-audio or VRAM footprint — all stated criteria here. A five-point Elo gap on one board at one refresh is also thin corroboration. That is a challenge, not a handover.

Also newly ingested: tencent/AuK is trending on Hugging Face as a text-to-speech pipeline with zero-shot cloning, editing and separation tags, but carries no independent quality or latency measurement yet, so it changes nothing today.

Breeze TTS 2 keeps the title on its verified package — sub-40ms first audio, 0.32 RTF on a warmed H100, ~7.7 GiB eager inference — with the standing caveat that the weights remain research/non-commercial. We will revisit as soon as StepAudio 2.5's weights and licence are pinned down or the board reading holds.

evidence: [1] [2]
13 Sep 2026
held

Weekly review: all 9 titles re-verified

All 9 picks were re-verified against the live leaderboards, repositories and release pages. 22 details moved with the evidence and the affected pick pages have been updated.

Under heightened review after this pass: GPT Image 2 (best image generation), Gemini Omni Flash (best video generation). A title only changes through full adjudication.

11 Sep 2026
challenged

Wan 3.0 extends its Artificial Analysis text-to-video lead over Gemini Omni Flash to five points

Artificial Analysis's text-to-video board on 11 September puts Alibaba Wan 3.0 first on 1242 Elo, with our pick Gemini Omni Flash second on 1237. Fal's post-trained Minimax H3 Max is third on 1232, MiniMax H3 Open Weights fourth on 1225 and Dreamina Seedance 2.0 720p fifth on 1220.

The useful detail is the trend rather than the single number. At our first read the two models were level on 1239; at the last read Wan 3.0 was one point ahead on 1240; it is now five points clear. That is the same board moving the same way three times, which is more than noise, and it is why the challenger note has been sharpened.

It is still not enough to move the title. This is one board measuring one of our four criteria. Gemini Omni Flash remains top of AA's no-audio text-to-video table and top on no-audio image-to-video, and there is still no independent read on Wan 3.0's character consistency or clip length — half our criteria are simply unmeasured for the challenger. A five-point Elo gap is also the sort of margin that sits inside a ranked range rather than establishing a clear win.

So: the pick holds, but on a narrower base than when we awarded it. Audio-inclusive generation is the weak flank, and if Wan 3.0 holds or extends this gap at the next read, or a second source corroborates it on consistency or clip length, the title should change hands.

evidence: [1] [2]
11 Sep 2026
challenged

GPT-Image-2.5-Sunburst challenges GPT Image 2 for best image generation

GPT-Image-2.5-Sunburst would take this title on today's evidence, but the case has not yet cleared our bar for a change, so it is recorded as the challenger for now. The adjudicator's reasoning:

GPT Image 2 has held this title since 21 April, but it is now behind on both boards we track for this slot.

Arena had already moved: GPT-Image-2.5-Sunburst leads text-to-image on 1421 against the incumbent's 1381, and — more important for the editability criterion — tops single-image edit on 1520 against 1461, with GPT-Image-2.5-Flare second on 1491. Today's Artificial Analysis fetch corroborates it. Its text-to-image arena now reads Flare (max) 1187, Sunburst (max) 1180, GPT Image 2 (high) 1171, dropping the incumbent to third on the board where it previously sat first. MAI-Image-2.6 is fourth on 1145 and Reve 2.1 fifth on 1127.

Two challengers split the boards. Flare is nominally ahead of Sunburst on Artificial Analysis, but by only 7 Elo, while Sunburst leads Arena on both text-to-image and single-image edit. Sunburst is therefore ahead on more of this slot's criteria, including editing, and takes the title; Flare becomes the challenger and could well overtake it as the AA sample grows.

Two caveats for buyers. Both 2.5 variants are recent arrivals on these boards, and the Flare/Sunburst margin on Artificial Analysis is well inside noise, so the ordering inside the 2.5 family is not settled. And cost remains a real consideration: Artificial Analysis prices GPT Image 2 at $211 per 1,000 images against $38.90 for MAI-Image-2.6, which still leads AA's image-editing arena on 1122. If you are editing at volume rather than chasing the top of the prompt-adherence charts, the cheaper Microsoft model remains the sane buy.

evidence: [1] [2]
10 Sep 2026
challenged

Wan 3.0 edges to nominal #1 on Artificial Analysis text-to-video, one Elo point clear

Artificial Analysis's text-to-video board now lists Alibaba Wan 3.0 at #1 on 1240 Elo, with Gemini Omni Flash second on 1239 — a one-point gap, where the two models were on an identical 1239 at the last read. Fal's post-trained Minimax H3 Max is third on 1235 and MiniMax H3 Open Weights fourth on 1228.

One Elo point on a single board is not a lead in any practical sense; it is a tiebreak placing that could reverse on the next refresh. It is also not corroborated anywhere else in the dossier: no second independent board puts Wan 3.0 ahead, and nothing here reads on clip length or character consistency, which are two of this slot's four criteria. Native audio remains the softest part of the pick's case — on AA's with-audio tables the pick has been level with or behind Wan 3.0 on text-to-video and behind fal's post-trained Minimax H3 Max on image-to-video — but that pressure is unchanged rather than newly decisive.

So the title holds and the challenger note is updated to record Wan 3.0's nominal top placing. Practical read for anyone choosing today: the top four on this board are close enough that prompt fit and workflow matter more than the ordering. If you need multi-turn conversational editing with audio in one pass, the pick is still the straightforward choice, with the caveat that Google's own known limitations list complete consistency across edits as unresolved. A second board putting Wan 3.0 clearly ahead, or a durable gap here, would take the title.

evidence: [1] [2]
10 Sep 2026
challenged

MiniCPM5-2B gets an official GGUF, tightening the challenge to Gemma 4 26B A4B

Gemma 4 26B A4B keeps the title, but one challenger has moved.

OpenBMB has now published an official GGUF repo for MiniCPM5-2B, tagged for tool-calling, long context and on-device/edge use (repo). That answers the main objection in our previous note, which was that MiniCPM5-2B had no quantised build at all and so no on-device story. It is still not enough to take the slot: the dossier shows no published file sizes for those GGUFs, no independent 4-bit quality-retention numbers and no tool-calling-under-quantisation measurements. On Artificial Analysis's small open-models index MiniCPM5-2B scores 15, below this pick's 17, so the case for it rests entirely on quality per GB — credible for a sub-4B model, but not something an independent board has yet measured.

The headline standings are unchanged. Qwen3.8 27B still leads aa-small-open at 34 (xhigh), with 28 medium, 26 low and 22 non-reasoning, against the pick's 17 at #10 (standings). Qwen3.6 27B (Reasoning) has entered the top five at 22. But Qwen3.8 27B remains disqualified on this slot's memory criterion: the only official quantised release is the ~31GB FP8 GPU repo, and community 4-bit GGUFs of a 27B dense model land near 17GB, above a 16GB machine. Qwen3.8-Flash-Next's NVFP4 build is GPU-oriented with no listed size and an "other" licence.

So the pick holds on deployability, not on raw score. A published, sized 4-bit build from either Qwen3.8 27B or MiniCPM5-2B, with tool-calling checked under quantisation, would likely settle this.

evidence: [1] [2] [3]
10 Sep 2026
changed

Claude Fable 5.1 replaces Claude Fable 5 as best overall LLM

The title moves within the family: Claude Fable 5.1 takes over from Claude Fable 5.

Artificial Analysis re-based its Intelligence Index twice this month (v4.2 on 4 September, v4.3 on 7 September). On the current v4.3 scale, Claude Fable 5.1 at max and xhigh effort shares first place on 53 with GPT-6 Astra (max/xhigh), with Claude Opus 5 on 51 and Fable 5 itself down on 50. The official Terminal-Bench board has moved from 2.1 to v4.0, and there Fable 5.1 leads at 57.9%, with Opus 5 on 51.8% and Fable 5 third at 44.5% — the old 83.8% figure we were quoting describes a retired board. Vellum's composite of quality, speed and cost (6 September) also puts Fable 5.1 first at 65%.

That leaves Arena text as the only board where Fable 5 still shows on top, at 1507 on the 2 September snapshot — but with a rank spread of 1-6 and Fable 5.1 (max) three points back at 1504, alongside Opus 4.6 (high) at 1505 and Opus 4.7 (high) at 1502. That is a tie band, not a lead worth holding a title on.

GPT-6 Astra is the honest alternative and stays as challenger: it ties Fable 5.1 on the Index at less than half the cost per Index task ($3.26 against $7.63), and leads Artificial Analysis's own Terminal-Bench v4.0 runs at 59.1% to 52.0%. But it has no Arena text placement at all and sits seventh on Vellum at 57.2%, so Fable 5.1 is ahead on more of this slot's criteria. Anyone paying by the token should still price Astra.

evidence: [1] [2] [3] [4] [5]
10 Sep 2026
changed

GLM-5.3 replaces Kimi K3 as best open-weight LLM

Kimi K3 held this slot on a one-point lead over GLM-5.3 on Artificial Analysis's Intelligence Index v4.2. That scale has been retired. On v4.3 the lead is gone: AA's leaderboard fetched on 9 September has GLM-5.3 (max) at 45 and Kimi K3 (max) at 44, while AA's open-weights page on 7 September names the two as level at 44 at the top of the open field. On the most favourable reading for the incumbent, quality is now a tie.

The slot's other criteria then decide it. The pick's own card already conceded that GLM-5.3 costs roughly $0.9 per Index task against K3's $2.3 and runs at about 84 tokens per second against K3's 41 — K3 is rack-scale to self-host and slow even on Moonshot's hosted API. Both models have weights genuinely released on Hugging Face (K3 on 27 July, GLM-5.3 on 28 August) and both ship under vendor-specific licences rather than a standard open one, so neither wins on licence. GLM-5.3 is therefore ahead or level on quality and clearly ahead on serving cost.

Two caveats. GLM-5.3 on Hugging Face is a text-generation model under Z.ai's custom glm-5.3 licence — check its terms before commercial use, and note it does not match K3's native vision, though vision is not a criterion here. GLM-5.3-Flash (MIT, 26 August) scores 42 and is the cheaper Flash-class option. DeepSeek-V4.1-Flash appeared on Hugging Face trending on 10 September under MIT with an FP8 checkpoint, but carries no independent score yet.

evidence: [1] [2] [3] [4] [5]
9 Sep 2026
challenged

Wan 3.0 takes nominal #1 on AA text-to-video, but on level points with Gemini Omni Flash

Artificial Analysis's text-to-video leaderboard now lists Alibaba Wan 3.0 first and Gemini Omni Flash second — but both are on 1239 Elo, so this is a tiebreak placing rather than a measured lead. Fal's post-trained Minimax H3 Max is four points back on 1235, MiniMax H3 Open Weights on 1228 and Dreamina Seedance 2.0 720p on 1222, which keeps the top of the board tightly bunched (Artificial Analysis).

We need more than a rank swap on equal points to move a title. Nothing in this week's evidence touches three of our four criteria: there is still no independent read on Wan's character consistency or single-pass clip length, and no fresh with-audio result to change the existing picture, where the pick is level with or behind Wan on text-to-video and behind Minimax H3 Max on image-to-video. Native audio remains the softest part of the incumbent's case, but it is not the part Wan has newly won.

So Gemini Omni Flash keeps the title, with the same caveats as before: 720p standard output with upscaling as a separate pass, so rivals' "15 seconds at 2K" figures are not like-for-like, and Google's own model card still lists complete consistency across edits as an open limitation.

What would settle it: a second independent board putting Wan 3.0 clearly ahead, or the same AA gap opening beyond noise and holding across successive reads — plus any credible third-party measurement of clip length and character consistency, where Wan is currently unmeasured.

evidence: [1] [2]
9 Sep 2026
changed

Breeze TTS 2 Open Weights takes best open-weight TTS from Fish Audio S2 Pro

Fish Audio S2 Pro has held this slot since March, but the Artificial Analysis provider-voice board no longer supports it. In the 9 September standings the pick sits at 1128 Elo in 25th place and is out of the top ten, while BreezeBlue's Breeze TTS 2 Open Weights is 7th at 1215 and the highest-placed open-weight model on the board. That is an 87-point gap on the slot's principal quality measure, not a rounding error, and the weights have been public on Hugging Face since 25 August 2026.

Breeze also wins the remaining criteria. BreezeBlue quotes under 40ms time-to-first-audio and 0.32 RTF on a warmed-up H100, against roughly 100ms first audio and 0.195 RTF on an H200 for S2 Pro, and it needs about 7.7 GiB for eager inference, so a 12GB card will do rather than H200-class kit. On licence there is no gain and no loss: Breeze's inference code is Apache 2.0 but the checkpoints carry a research-and-non-commercial licence, exactly as restrictive as the Fish Audio Research License. Commercial deployment still needs a separate agreement in both cases.

Two caveats worth weighing before you swap. Breeze covers English and Chinese only, where S2 Pro spans 80-plus languages; if you need broad language coverage the old pick remains the better tool despite the Elo gap. And Step-Audio-EditX, at 1104 Elo, is behind both but keeps the loosest code licensing, so it stays the fallback for self-hosters who must ship commercially.

evidence: [1] [2] [3]
9 Sep 2026
challenged

Breeze TTS 2 challenges Fish Audio S2 Pro for best open-weight TTS

Breeze TTS 2 would take this title on today's evidence, but the case has not yet cleared our bar for a change, so it is recorded as the challenger for now. The adjudicator's reasoning:

The Artificial Analysis provider-voice arena has settled the question. BreezeBlue Breeze TTS 2 Open Weights now sits at #7 overall on 1215 Elo — the highest-placed open-weights model on the board — while Fish Audio S2 Pro has fallen out of the top ten altogether. That is not a one-off reading: the same board had Breeze at 1,220 against 1,125 for S2 Pro shortly after the weights went up on Hugging Face on 25 August 2026, and the gap has held through today's fetch.

Breeze also takes the two engineering criteria. BreezeBlue quotes under 40ms time-to-first-audio on an H100 with roughly 7.7GiB for eager inference, a 12GB GPU minimum and 24GB for the fast path; S2 Pro's quoted ~100ms first audio and 0.195 RTF assume an H200-class card. On licence there is nothing to choose between them: Breeze ships Apache-2.0 inference code but puts weights, derivatives and self-hosted outputs under the BreezeBlue Research and Non-Commercial Licence, needing written authorisation from RESONIA — the same blocker as the Fish Audio Research License.

Two caveats stay on the record. Breeze is English and Chinese only, a sharp narrowing from S2 Pro's 80+ languages; if you need broad multilingual coverage, the old pick is still the one to reach for. And if you need to ship commercially from self-hosted weights, neither of these helps you — Step-Audio-EditX, at 1104 Elo and Apache-2.0 on the repository code (though with no licence file on the weights repo), remains the least restrictive on paper and runs in 12GB of VRAM.

evidence: [1] [2]
9 Sep 2026
challenged

Qwen3.8 27B's lead widens on Artificial Analysis, but its 16GB fit is still unverified

Artificial Analysis has re-scored its small open-models board again. As of 9 September, Qwen3.8 27B holds the top three slots at 34 (xhigh), 31 (medium) and 29 (low), with its non-reasoning mode at 22. Gemma 4 26B A4B, this slot's title holder, sits at #10 on 17. Both sets of numbers have come down from the August scoring (41/35/34 against 26), so the relative picture is unchanged: Qwen3.8 27B is still the quality leader in this size class, and the gap in index points has if anything widened slightly.

That still does not take the title. This slot is judged on quality per GB, tool-calling under quantisation and running in 16GB of RAM, and Qwen3.8 27B remains without an official quantised release with published file sizes, without independent 4-bit quality-retention figures and without tool-calling-under-quantisation numbers. Its fit on a 16GB machine is therefore unproven, which is a stated criterion rather than a nice-to-have. Gemma 4 26B A4B keeps the title on deployability: Apache-2.0 weights, a 14.2GB dynamic UD-Q4_K_XL QAT build, native function calling.

The rest of this week's dossier is noise for this slot — two unrelated arXiv method papers on preference-data filtering and phase geometry in transformers.

The title stays under review. A verified quantised Qwen3.8 27B build with published sizes and any independent quantised tool-calling evidence would very likely take the slot on the next pass.

evidence: [1] [2] [3] [4]
9 Sep 2026
challenged

GLM-5.3 challenges Kimi K3 for best open-weight LLM

GLM-5.3 would take this title on today's evidence, but the case has not yet cleared our bar for a change, so it is recorded as the challenger for now. The adjudicator's reasoning:

Artificial Analysis now lists GLM-5.3 (max) at 45 on its Intelligence Index, one point ahead of Kimi K3 (max) at 44 — the reverse of the 50-vs-49 reading that put K3 in this slot in July (AA models).

Treat the quality gap as a tie: one point, in opposite directions across the two readings, is inside the noise. What settles it is the rest of the slot's criteria. GLM-5.3's full weights went public on Hugging Face between 25 and 28 August, so the availability objection that kept it as a challenger is gone. On cost and speed it is not close: roughly $0.9 per Index task against K3's $2.3, and about 84 tokens per second against K3's 41. K3 keeps the bigger context window and native vision, and it remains a serious model, but paying more than twice as much for half the throughput to score level is no longer defensible.

Licence is a wash rather than a win. GLM-5.3 is tagged 'other' on Hugging Face, not MIT; K3 shipped under Moonshot's own custom terms with attribution and MaaS conditions attached. Both need reading before commercial deployment.

The other open-weight contenders stay behind. DeepSeek-V4-Flash-Vision-Exp scores 42 and is an explicitly experimental Flash-class checkpoint, though its ~168GB FP8/FP4 footprint makes it much the easiest to self-host. GLM-5.3-Flash also sits at 42, MIT-licensed, and is the sensible pick if you need something small and permissive.

Self-hosting GLM-5.3 is still rack-scale work: reported local serving needs 8x H200 or 10-12x H100 at FP8.

evidence: [1]
9 Sep 2026
challenged

Claude Opus 5 logged as challenger: leads the independent SWE-bench Verified aggregate, where GPT-6 Astra has no run

GPT-6 Astra keeps the title, but with a named challenger.

The new evidence is benchlm's independent SWE-bench Verified aggregate, refreshed 9 September, which has Claude Opus 5 first at 96%, ahead of Claude Mythos 5 (95.5%) and Claude Fable 5 (95%). GPT-6 Astra does not appear in that top ten at all — it has no independent SWE-bench Verified run, which was already flagged in the pick's rationale.

That is not enough to move the title. Where the two models are directly compared, Astra is still ahead: it is first on the official Terminal-Bench board (v4.0, 58.2%) against Opus 5's 51.8% in third, and its max tier is first on Arena WebDev at 1796 against claude-opus-5-max's 1688. Those two boards carry the slot's agentic and tool-calling criteria between them. A single board on which the pick is simply absent is a gap in the record, not a defeat.

It is worth logging as a challenge rather than ignoring, because SWE-bench Verified is the closest thing on our boards to repo-scale performance, and it is the one criterion where we have no reading for the pick at all. Two things would settle it: an independent SWE-bench Verified run for Astra, or a second board putting Opus 5 ahead on agentic or tool-calling work. Note also that vals.ai removed SWE-bench Verified from its coding index in August as saturated, so a 96% v 97% ordering there should not be over-read either way.

Claude Fable 5.1 sits second on both Terminal-Bench 4.0 (57.9%) and Arena WebDev (1764) and remains close behind.

evidence: [1] [2]
9 Sep 2026
changed

GPT-6 Astra takes best LLM for coding from Claude Opus 5

Two independent boards now point the same way. The official Terminal-Bench board — now on version 4.0, a harder task set than the retired 2.1 — has GPT-6 Astra first at 58.2%, with Fable 5.1 second (57.9%) and Opus 5 third at 51.8% — the agentic-coding criterion this slot leans on hardest. Arena WebDev has gpt-6-astra-max first at 1796, ahead of claude-fable-5.1-max (1764) and claude-opus-5-max (1688), which is where Opus 5 now sits at #3. The Terminal-Bench result is not a one-off: vals.ai's own Terminal-Bench 2.1 table also has GPT-6 Astra on top (87.27%, ahead of GPT-5.6 Sol at 85.77% and Claude Fable 5.1 at 85.02%), and it leads vals' Code Migration table too. The reading has held across successive fetches through early September.

Where the two Astra variants split the boards, Astra leads the agentic and repo-scale tables while the max tier tops human-preference WebDev, so the title goes to Astra.

What Opus 5 keeps is SWE-bench Verified: it is still #1 on the independent aggregate at 96%. But vals has dropped that benchmark from its coding index as saturated, and the top three there are within a point of each other. Astra has no independent SWE-bench Verified run yet, so its repo-scale case rests on Code Migration alone — the main reason to hold this at medium confidence rather than high.

Practical note: Astra costs $10/$50 per million tokens against Opus 5's $5/$25. For cheap, high-volume repo-scale batch work, Opus 5 remains the better buy.

evidence: [1] [2]
9 Sep 2026
challenged

GPT-6 Astra challenges Claude Opus 5 for best LLM for coding

GPT-6 Astra would take this title on today's evidence, but the case has not yet cleared our bar for a change, so it is recorded as the challenger for now. The adjudicator's reasoning:

The title moves to GPT-6 Astra Max. On the agentic side, vals.ai's Terminal-Bench 2.1 now has Astra first at 87.27%, ahead of GPT-5.6 Sol (85.77%), Claude Fable 5.1 (85.02%) and Opus 5 in fourth at 84.64% — and Opus 5's own figure falls to 81.27% once its server-side fallbacks to Opus 4.8 on refused tasks are counted as failures. Astra also tops vals' Code Migration table, which is the nearest thing on that board to repo-scale work. That is not a one-board reading: Arena WebDev has had Astra at #1 across both fetches in this cycle, 1,796 on 9 September against claude-fable-5.1-max at 1,764 and claude-opus-5-max at 1,688, with our outgoing pick third.

Opus 5 keeps #1 on the independent SWE-bench Verified aggregate at 96%, but the top three there sit within a point of each other and vals has dropped that benchmark from its coding index as saturated, so it no longer carries the weight it did when Opus 5 took the title in July.

Claude Fable 5.1 Max is the other contender and leads the Vals Index overall, but it trails Astra on both boards this slot tracks and its refusals on bio- and cyber-adjacent tasks remain an operational cost. Astra is dearer at $10/$50 per million tokens against Opus 5's $5/$25, and there is still no independent SWE-bench Verified run for it — that is the main reason for medium rather than high confidence. Teams already happy with Opus 5's cost profile have no urgent reason to switch; new agentic setups should default to Astra.

evidence: [1] [2] [3]
9 Sep 2026
challenged

GPT-Image-2.5-Sunburst tops both Arena boards, putting GPT Image 2's title under challenge

GPT Image 2 keeps the title this week, but its position has weakened materially. OpenAI has shipped ChatGPT Images 2.5, and Arena now ranks gpt-image-2 (medium) third on both boards that matter here: third in text-to-image on 1381, behind GPT-Image-2.5-Sunburst (1421) and Flare (1399), and third in single-image edit on 1461, behind Sunburst (1520) and Flare (1491). Those are the slot's two principal criteria, and the margins — 40 and 59 Elo — are wider than measurement noise.

What stops a change today is corroboration. The lead rests on a single board, read once. Artificial Analysis has not yet rated any GPT-Image-2.5 variant: its text-to-image arena still has GPT Image 2 (high) first on 1178, 29 clear of MAI-Image-2.6 (1149) and 51 clear of Reve 2.1 (1127). There is also no verified official model page for GPT-Image-2.5 in front of us, and we will not point a title at an unconfirmed listing.

The previous challenger, MAI-Image-2.6, is now the lesser story. It still leads Artificial Analysis's editing arena on 1122 against GPT Image 2 (high) on 1117 — a five-point gap, effectively level — and remains far cheaper at $38.90 per 1,000 images against $211. That is a cost argument, not a quality one.

If Artificial Analysis rates the 2.5 models, or Arena holds these standings on a second reading, this title moves to Sunburst.

evidence: [1] [2] [3]
9 Sep 2026
held

Weekly review: all 4 titles re-verified

All 4 picks were re-verified against the live leaderboards, repositories and release pages. 10 details moved with the evidence and the affected pick pages have been updated.

Under heightened review after this pass: Claude Opus 5 (best llm for coding), Kimi K3 (best open-weight llm), Claude Fable 5 (best overall llm). A title only changes through full adjudication.

9 Sep 2026
held

Weekly review: all 6 titles re-verified

All 6 picks were re-verified against the live leaderboards, repositories and release pages. 11 details moved with the evidence and the affected pick pages have been updated.

Under heightened review after this pass: GPT Image 2 (best image generation), Fish Audio S2 Pro (best open-weight tts). A title only changes through full adjudication.

8 Sep 2026
challenged

MiniCPM5-2B logged as a challenger; AA small-open board re-scored again

No change to the title. Gemma 4 26B A4B stays because nothing here disturbs the argument that won it: a 14.2GB dynamic 4-bit build that actually loads in 16GB, with native function calling.

Two things worth recording. First, Artificial Analysis has re-scored the small open board again: Qwen3.8 27B now reads 34 at xhigh effort, 31 medium, 29 low and 22 non-reasoning, down from 52/44/43 in the 22 August snapshot (AA small open). It remains the quality leader in the class, and the underlying blocker is unchanged — no official quantised release with published file sizes, so its fit in 16GB is still unproven. A new entrant, Qwen3.6 35B A3B (Reasoning), arrives at #5 with 22. The MoE shape is interesting for on-device decode speed, but 35B total gives it no obvious path into 16GB at 4-bit and there is no published sizing to check.

Second, MiniCPM5-2B is trending on Hugging Face, with the card tagging long context, tool-calling and edge deployment (model card). At 2B it would be a quality-per-GB argument rather than a quality one, and right now there is no independent benchmark placement and no quantised file sizes to verify. Logged as a challenger, not a contender.

evidence: [1] [2]
8 Sep 2026
challenged

GPT-6 Astra draws level with Claude Fable 5.1 on the Artificial Analysis Intelligence Index

Artificial Analysis has re-scored its Intelligence Index leaderboard and the top of the board is now effectively level. The listed #1 is Claude Fable 5.1 (Adaptive Reasoning, Xhigh Effort, Default Fallback) on 53, with GPT-6 Astra (max) at #2 and GPT-6 Astra (xhigh) at #3, both also on 53. Claude Fable 5.1 (High Effort) sits fourth on 51, with GPT-6 Astra (high) fifth on the same score.

Two things follow. First, the incumbent keeps the slot: on the one independent composite that moved, the Fable line is still ranked first, and nothing in this evidence touches the Arena text board or the official Terminal-Bench 2.1 board, where Fable 5 holds 83.8% with Claude Code at xhigh effort.

Second, the GPT-6 Astra challenger note is now out of date. It described Astra as sitting two to three points behind on the Index; on the current snapshot it is tied on score and separated only by ordering. That is a genuine narrowing, but a tie on a single composite is not a win on the slot's criteria, and Astra still has no Arena or Terminal-Bench 2.1 placement to corroborate it.

Note also that Claude Opus 5, previously the strongest case for a change on this board, is not visible in the current top five. Index scores have clearly been rebased — the absolute numbers here are lower than the 62/63 figures carried in the caveats — so read exact standings off the live board rather than comparing across snapshots.

evidence: [1]
7 Sep 2026
challenged

Microsoft's VibeVoice-ASR-Streaming-7B logged as an unverified challenger; ARK-ASR-3B holds

No change this week. The bulk of the incoming evidence is literature rather than evaluation: papers on ASR hallucination under environmental degradation, a word-level backdoor attack (GhostWord), turn-aware streaming supervision, a trilingual spoken-dialogue fact-checking benchmark, and a Mandarin-English code-switching dataset. Interesting reading, but none of it scores a model against the criteria this slot is judged on, so none of it moves ARK-ASR-3B.

The one item worth recording is microsoft/VibeVoice-ASR-Streaming-7B, which appeared on Hugging Face's trending list on 2 September 2026. From the listing we can see it is an automatic-speech-recognition pipeline tagged for streaming and for four languages (en, zh, es, pt). That is the whole of the public signal. There is no independent leaderboard result, no reported throughput, and the listing surfaces no licence — and without a licence you cannot even confirm it belongs in an open-weight slot, let alone whether it beats a 4.76 mean WER at RTFx 490.98.

A streaming 7B from Microsoft is plausibly relevant to voice-agent workloads where latency, not batch WER, decides the design. But this slot ranks on independent WER first, and until the model is scored with the Open ASR Leaderboard's own harness there is nothing to compare. It joins the existing challenger list on the same terms as Orze-ASR-3Way: a repo exists, a claim does not. Check the card yourself for licence terms before building against it.

evidence: [1] [2] [3] [4] [5]
7 Sep 2026
held

Weekly sweep: 10 slots reviewed, no changes

75 candidate event(s) examined across all slots; none met promotion criteria. All last-reviewed dates refreshed.

6 Sep 2026
challenged

Qwen3.5 Omni enters the Artificial Analysis STT top five; ARK-ASR-3B holds

Artificial Analysis' speech-to-text board reshuffled on 6 September 2026, with Qwen3.5 Omni Flash listed at #1 and Qwen3.5 Omni Plus at #2, ahead of Nova 2 Pro, Amazon Transcribe and Universal-3 Pro. Three of those five are hosted commercial services, so they are out of scope for this slot; the two Qwen3.5 Omni entries are the only plausible open-weight challengers in the listing.

They do not clear the bar. The listing gives us a rank and a percentage and nothing else: no licence, no confirmation that weights are downloadable, no throughput figure and no language coverage. The percentages as recorded are also hard to read as error rates — the #1 entry is shown at 13.5% while the #5 entry is shown at 3.1% — so whatever the board is ordering on, it is not a straight word error rate, and we will not restate those numbers as WER.

Crucially, this slot is judged on the Open ASR Leaderboard harness, where ARK-ASR-3B still holds at 4.76 mean WER with RTFx 490.98 under Apache-2.0. Nothing in today's evidence scores a Qwen3.5 Omni variant on that harness, so there is no like-for-like comparison to act on.

ARK-ASR-3B therefore keeps the title. Qwen3.5 Omni goes on the challenger list pending three things: confirmation that the weights are published under a usable licence, an Open ASR Leaderboard entry run with the leaderboard's own scorer, and a throughput number. If those land and beat 4.76, this becomes a change.

evidence: [1] [2]
6 Sep 2026
challenged

NVIDIA publishes an NVFP4 build of Qwen3.8-Flash-Next; Gemma 4 26B A4B holds the slot

The only new item this cycle is an NVIDIA-published NVFP4 quantisation of Qwen3.8-Flash-Next, trending on Hugging Face and tagged image-text-to-text, quantized, FP4 against the Qwen/Qwen3.8-Flash-Next base model. That is worth noting because the standing objection to Flash-Next in this slot has been the absence of an official quantised release, but it does not resolve it. NVFP4 via NVIDIA's Model Optimizer is a datacentre GPU format, not a laptop-friendly GGUF, and the listing carries no published file size, so the 16GB fit remains untested. There is still no independent evaluation of Flash-Next at any precision, and nothing at all on tool-calling retention once quantised — the two criteria that decide this slot.

The other item in the dossier is the 22 August Artificial Analysis open small-model board, which is already reflected in the current caveats: Qwen3.8 27B leads on intelligence, with Gemma 4 26B A4B the board pick on deployability grounds. Nothing there is new.

So the title stays with Gemma 4 26B A4B. The argument for it is unchanged and narrow: a 14.2GB QAT int4 GGUF that actually loads on a 16GB machine, Apache-2.0 weights, and native function calling. The moment someone publishes an independently measured 4-bit build of a Qwen3.8-class model that fits the same envelope and holds up on tool calls, this slot should turn over. A vendor FP4 upload is not that.

evidence: [1] [2]
6 Sep 2026
challenged

GPT-6 Astra gets its first independent placement — second, behind the Claude Fable line

GPT-6 Astra is no longer a vendor-claims-only entry. Artificial Analysis today lists it at #2 on the Intelligence Index (max effort, 55) and #3 (xhigh, 54). That answers the main objection in its previous challenger note, but it places OpenAI's frontier model behind the incumbent rather than ahead of it: the same board now has Claude Fable 5.1 (adaptive reasoning, max effort) at #1 with 57.

The other movement is within the Anthropic stack. Claude Opus 5, which had been the strongest case for a change after taking the Intelligence Index lead, has dropped to #4 (54) and #5 (53) on Artificial Analysis, and to #3 on Vellum's board at 64.7%. Vellum now shows Claude Fable 5.1 first at 65%, with Claude Mythos 5.1 alongside it at 65%. On both independent aggregators, the Claude Fable line has reclaimed the top spot it lost in the summer.

Two caveats. First, the boards are now listing a 5.1 point release while this slot's pick is recorded as Claude Fable 5; a same-line version bump is not a title change, but readers should expect the card's benchmark figures — the Arena text lead and the 83.8% on the official Terminal-Bench 2.1 board — to refer to the earlier build until fresh placements appear. Second, the margins are small: 57 against 55, and 65% against 64.7%. Treat this as the incumbent holding on rather than pulling away, and read the exact standings off the live boards.

evidence: [1] [2]
6 Sep 2026
challenged

gpt-6-astra-max takes #1 on Arena WebDev, but no agentic evidence yet

A new model, gpt-6-astra-max, has entered the Arena WebDev board at #1 with 1797, ahead of claude-fable-5.1-max at 1762, claude-opus-5-max at 1688, qwen3.8-max-0902 at 1686 and kimi-k3-max at 1674. That is a 35-point margin over the previous leader and a clear gap over the Opus line — larger than the four-point shuffles we have logged on this board over the past few weeks, so it is worth registering rather than filing as noise.

It does not move the title. Arena WebDev is a human-preference board scoring web front-end output; none of this slot's criteria — agentic coding benchmarks, repo-scale task performance, tool-calling reliability — are measured by it. We have seen the same pattern twice already: claude-fable-5.1-max and qwen3.8-max-0902 both topped this board without a single independent agentic run appearing afterwards, and neither displaced Opus 5.

What would change our mind is an independent SWE-bench Verified or Terminal-Bench 2.1 run on a third-party harness. Opus 5 holds the pick on vals.ai's SWE-bench Verified board at 97.0% and sits second on their Terminal-Bench 2.1 table at 84.64%. Those numbers are already under pressure — GPT-5.6 Sol leads Terminal-Bench 2.1 at 85.77%, DeepSeek V4 Pro is within 0.6 points on SWE-bench Verified, and Opus 5's terminal figure drops to 81.27% if its server-side fallbacks to Opus 4.8 count as failures. A genuinely strong agentic showing from the new model would likely settle it. Until one is published, Opus 5 stays.

evidence: [1]
6 Sep 2026
held

Weekly review: all 8 titles re-verified

All 8 picks were re-verified against the live leaderboards, repositories and release pages. 17 details moved with the evidence and the affected pick pages have been updated.

Under heightened review after this pass: Claude Opus 5 (best llm for coding), Fish Audio S2 Pro (best open-weight tts). A title only changes through full adjudication.

5 Sep 2026
challenged

Qwen3.8 27B still leads the small-model board, but its Artificial Analysis scores have been marked down

Artificial Analysis's small open-source board has been re-scored, and the numbers on this slot's challenger card no longer match it. Qwen3.8 27B now sits at 41 at xhigh effort, 35 at medium and 34 at low, against 52/44/43 when we last recorded the board on 22 August. Qwen3.6 27B (reasoning) has likewise dropped from 38 to 29. A new entry, Qwen3.8 27B in non-reasoning mode, comes in at #5 with 26 — level with Gemma 4 26B A4B.

The ordering is unchanged: Qwen3.8 27B is still the quality leader in this size class, and still by a clear margin at its higher effort settings. What has changed is the size of that margin, which is now roughly 15 points rather than 26, and the fact that the cheapest configuration of the challenger scores no better than the pick.

None of this moves the title, because none of it touches the criteria. There is still no official quantised Qwen3.8 27B release with published file sizes, no independent measurement of quality retention at 4-bit, and no tool-calling numbers under quantisation. Until someone publishes a build that demonstrably loads and behaves in 16GB, the 27B dense models remain unproven on the one axis this slot cares about.

Also in this window: SGLang v0.5.19 added support for Qwen3.8 2.4T-A95B. That is a datacentre model and has no bearing here; we note it only because it will show up in searches for the same family name.

Gemma 4 26B A4B holds, on its 14.2GB QAT int4 build.

evidence: [1] [2]
5 Sep 2026
challenged

GPT-6 Astra formally launched, but no independent numbers yet

OpenAI has published a launch post for GPT-6 Astra, describing it as its "most intelligent and aligned model yet" with state-of-the-art capabilities across computer use, coding, cybersecurity and science (openai.com). That settles the earlier uncertainty over whether the model was real and public: it is now an announced release rather than a rumour circulating on social posts.

What it does not do is move the title. The only evidence we have is the vendor's own announcement. There is no Arena text placement, no Artificial Analysis Intelligence Index figure, and no appearance on the official Terminal-Bench 2.1 board — the three independent sources this slot leans on. On our stated criteria, reasoning quality, multimodal capability and long-context reliability are all asserted by OpenAI rather than measured by anyone else. Launch-post benchmark claims are exactly the category of evidence that cannot carry a title here, however plausible they look.

So Claude Fable 5 holds. It remains top of the Arena text board and first on the official Terminal-Bench 2.1 board at 83.8% with Claude Code at xhigh effort. The standing caveats also still apply: the Arena top band is tight, and Claude Opus 5 leads Artificial Analysis's Intelligence Index at 63 to Fable's 62 at roughly a quarter less per Index task, which remains the strongest live case against the incumbent.

We expect third-party placements for Astra within days to weeks. If it lands at the top of Arena text or Terminal-Bench 2.1, or takes the Intelligence Index outright, this slot will change quickly. Until independent numbers exist, it stays a challenger.

The vLLM release candidate in this batch is routine and has no bearing on the slot.

evidence: [1]
3 Sep 2026
challenged

Wan 3.0 still nominally top on AA text-to-video, but the gap has closed to zero

Artificial Analysis's text-to-video leaderboard updated on 3 September with Alibaba's Wan 3.0 listed first on 1,238 and Gemini Omni Flash second on 1,238 — the same score. That is a narrowing, not a gain: on 20 August the same board had Wan 3.0 on 1,247 to the pick's 1,239, an eight-point lead. Fal's post-trained Minimax H3 Max is third on 1,235, MiniMax H3 Open Weights fourth on 1,227 and ByteDance Seed's Dreamina Seedance 2.0 720p fifth on 1,221, so the top five are inside eighteen points and the ordering at the head of the table is effectively a coin toss.

On the slot's criteria this changes nothing. A tied Elo on one sub-board is not evidence that Wan beats the pick on visual quality, and there is still no independent read on Wan's character consistency, native audio behaviour or maximum clip length — the three places where Gemini Omni Flash's multi-turn editing and synced audio on every clip have been the differentiator. The pick also continues to lead AA's no-audio text-to-video board and arena.ai's Text-to-Video Arena.

Wan 3.0 stays the lead challenger and remains the most likely candidate to take this title, but it needs either a durable lead on the with-audio board or a second, independent evaluation covering consistency and clip length. Title holds.

evidence: [1]
3 Sep 2026
challenged

Claude Fable 5.1 Max takes Arena WebDev #1 by a wide margin; logged as challenger

Anthropic's claude-fable-5.1-max has entered the Arena WebDev board straight at #1 with 1765 Elo, ahead of qwen3.8-max-0902 (1688), claude-opus-5-max (1687), kimi-k3-max (1674) and qwen3.8-max (1669). Unlike the Qwen3.8-Max result we logged previously — a four-point gap that was a tie in practice — this is a 77-point lead over the next entry and a 78-point lead over the incumbent's Arena variant, which is not noise.

It is still not enough to move the title. Arena WebDev is a human-preference board for web front-end work; it does not measure any of this slot's stated criteria — agentic coding benchmarks, repo-scale task performance, or tool-calling reliability. Fable 5.1 has no independent SWE-bench Verified or Terminal-Bench 2.1 run in the evidence to date, and the earlier Fable 5 terminal figures already came with a disclosed harness caveat (Vals reports Opus 4.8 used as a refusal fallback in both Fable 5's and Opus 5's runs).

Claude Opus 5 therefore holds, on the strength of its independent SWE-bench Verified result and its price position. The existing caveats stand: GPT-5.6 Sol leads Opus 5 on vals.ai's Terminal-Bench 2.1 table, and DeepSeek V4 Pro is within a point on SWE-bench Verified as an open-weight option. We will revisit if an independent agentic or repo-scale run for Fable 5.1 appears.

Also noted but not decision-relevant: PaperCompiler, a new arXiv method for repository-level paper-to-code generation, which reports no model ranking bearing on this slot.

evidence: [1] [2]
2 Sep 2026
challenged

Community mixed-precision GGUFs appear for Qwen3.8 27B, but the 16GB question is still open

The gap that has kept Qwen3.8 27B in the challenger column is starting to close, but it has not closed. ISTA-DASLab has published a mixed-precision GGUF conversion of Qwen3.8 27B (tagged GSQ/RCO, multimodal, image-text-to-text) which is now trending on Hugging Face (model card). That is the first quantised build of the quality leader in this size class to gain visible traction.

It does not yet move the title. This slot is decided on quality per GB, tool-calling under quantisation, and whether the thing loads in 16GB of RAM — and none of those are established for this build in the evidence to hand. There are no published file sizes, so the 16GB fit remains unproven; there is no independent measurement of quality retention against the bfloat16 parent, which is precisely the failure mode that made Gemma 4's QAT int4 build the pick (Unsloth measured 70.2% top-1 for a naive Q4_0 conversion against 85.6% for a dynamic build); and there is no report on function calling at reduced precision.

The leaderboard picture is unchanged. Artificial Analysis's small open-source board still shows Qwen3.8 27B holding the top three positions at 52/44/43 against this pick's 26, with Qwen3.6 27B at 38 and Muse Glimmer at 35 (board).

So Gemma 4 26B A4B holds on deployability, not on score. If someone publishes sizes and a quality-retention or tool-calling check for one of these Qwen3.8 quants, expect this title to change.

evidence: [1] [2] [3]
2 Sep 2026
challenged

Qwen3.8-Max-0902 takes #1 on Arena WebDev, joining the challenger list behind Claude Opus 5

Alibaba's new checkpoint, qwen3.8-max-0902, has moved to #1 on the Arena WebDev coding board at 1,691 Elo, displacing claude-opus-5-max, which now sits second at 1,687. The rest of the top five is kimi-k3-max at 1,674, the earlier qwen3.8-max at 1,669 and claude-opus-5-high at 1,661 (Arena WebDev).

A 4-point Elo gap is not a result. On a board of this size that margin is comfortably inside the usual confidence interval, and the two models are best read as tied. More importantly, Arena WebDev measures human preference between generated web front-ends. It tells you little about the three things this slot is judged on: agentic coding benchmarks, repo-scale task completion and tool-calling reliability. Kimi K3 has been sitting in the same neighbourhood on this board for weeks without shifting the title, for the same reason.

So Qwen3.8-Max goes on the challenger list rather than into the pick. What would move it: an independent agentic run — SWE-bench Verified or Terminal-Bench 2.1 on a third-party harness such as vals.ai — showing it at or above Opus 5's 97.0% and 84.64%. Nothing in this cycle provides that.

The existing caveats stand unchanged. Opus 5's terminal lead remains genuinely contested against GPT-5.6 Sol, and its Terminal-Bench figure still depends on how you score server-side fallbacks to Opus 4.8. Claude Opus 5 holds the title.

evidence: [1]
31 Aug 2026
challenged

Wan 3.0 stretches its with-audio lead over Gemini Omni Flash to eight points

Artificial Analysis's text-to-video board now puts Alibaba Wan 3.0 first on 1247 with Gemini Omni Flash second on 1239, with MiniMax H3 Open Weights third on 1228 and ByteDance's Dreamina Seedance 2.0 720p fourth on 1221. That widens Wan's margin over the pick from four points to eight on the same board, and it is the second consecutive read in which Wan sits ahead rather than a one-off.

It is still not enough to move the title. Eight Elo points on a single crowd-vote sub-board, on a table where AA had lately been publishing a ranked range spanning several positions for Wan, is not a decisive result on visual quality, and the slot is judged on four criteria rather than one. Gemini Omni Flash continues to lead AA's no-audio text-to-video table and its image-to-video board without audio, and there is still no independent measurement of Wan 3.0's clip length or character consistency to set against the pick's multi-turn editing through the Interactions API, which is the practical reason it holds the slot.

The only other item in this cycle is a community-converted MiniMax-H3-experimental repo trending on Hugging Face. That is a third-party packaging of a model already tracked as a challenger and carries no evaluation of its own, so it does not shift anything.

If Wan 3.0's lead holds or grows once AA's sample count closes on the pick's, and any independent read appears on length or consistency, this becomes a change rather than a challenge.

evidence: [1] [2] [3]
31 Aug 2026
challenged

DeepSeek-V4-Flash-Vision-Exp lands under MIT; Kimi K3 keeps the open-weight title

An experimental member of the DeepSeek V4 family, deepseek-ai/DeepSeek-V4-Flash-Vision-Exp, appeared on the Hugging Face trending list on 31 August with safetensors weights, fp8 and 8-bit variants, an image-text-to-text pipeline, eval-results tags and — notably — an MIT licence (model card). On licence and "weights actually released", that is stronger than the incumbent: Kimi K3 ships under Moonshot's own custom terms with attribution and MaaS conditions above certain thresholds.

It does not take the title. This is a Flash-class, explicitly experimental checkpoint, and nothing in the current evidence gives it an independent head-to-head score against K3 or against the GLM-5.3 weights already logged as a challenger. Flash-tier releases from every lab so far trade quality for serving cost, and serving cost is only one of four criteria here. We would want an Artificial Analysis-style Intelligence Index placement, or another third-party evaluation, on the downloadable weights before treating a V4 checkpoint as the frontier open-weight model.

Also in the dossier and not moving the needle: orcarouter/Qwen3.8-Flash-Next-Uncensored-GGUF, a community abliterated quantisation of an existing Flash-class Qwen rather than a new model (listing).

Kimi K3 therefore holds the slot, still on the strength of being the highest-scoring thing you can download, and still with the same caveats: enormous to self-host and slow in practice. The interesting question for the next few weeks is whether a full, non-experimental DeepSeek V4 arrives under the same MIT terms.

evidence: [1] [2]
31 Aug 2026
held

Weekly sweep: 10 slots reviewed, no changes

49 candidate event(s) examined across all slots; none met promotion criteria. All last-reviewed dates refreshed.

30 Aug 2026
held

Weekly review: all 10 titles re-verified

All 10 picks were re-verified against the live leaderboards, repositories and release pages. 25 details moved with the evidence and the affected pick pages have been updated.

Under heightened review after this pass: Kimi K3 (best open-weight llm), Fish Audio S2 Pro (best open-weight tts). A title only changes through full adjudication.

28 Aug 2026
challenged

Breeze TTS 2 Open Weights enters the Artificial Analysis top five, challenging Fish Audio S2 Pro

The Artificial Analysis TTS leaderboard now lists BreezeBlue Breeze TTS 2 Open Weights at #5 with an Elo of 1,217, behind Cartesia Sonic 3.6 (1,285), SpeechifyAI Simba 3.2 (1,241), Alibaba Qwen-Audio-3.0-TTS-Plus (1,240) and VUI Labs Luna TTS (1,225). Fish Audio S2 Pro, which holds this title partly on the strength of being the highest-ranked open-weights entry on that same board, does not appear in the published top five. On the face of it, an open-weights rival has overtaken the current pick on the arena metric we track.

That is enough to log a challenger, not to change the title. The leaderboard row tells us nothing about Breeze TTS 2's actual licence terms — "Open Weights" in a model name is not a verified licence, and this slot cares specifically about whether commercial self-hosting is permitted. We also have no time-to-first-audio figures, no parameter count or VRAM requirement, and no independent confirmation that checkpoints are downloadable and reproduce the arena score. Fish S2 Pro's current Elo is not stated either, so the size of the gap is inferred from its absence from the top five rather than measured.

What would settle it: a published licence file, a reproducible local inference path with latency and GPU figures, and a leaderboard snapshot showing both models' scores side by side. Until then Fish Audio S2 Pro keeps the slot, with the standing caveat that its licence remains non-commercial and Step-Audio-EditX stays the Apache-2.0 fallback. The two arXiv items in this cycle — a Sanskrit chant TTS pipeline and the SPAR-K early-exit decoding scheme — do not bear on the title.

evidence: [1] [2] [3] [4]
28 Aug 2026
challenged

GLM-5.3 weights land on Hugging Face, drawing level with Kimi K3 but not past it

Z.ai's full GLM-5.3 is now downloadable: the zai-org/GLM-5.3 repository appeared on Hugging Face's trending list on 25 August, and its public release was confirmed on 28 August alongside practical serving notes — roughly 10–12x H100 (or 8x H200) for FP8, about 390–430GB at 4-bit/NVFP4, and 230–250GB under aggressive 2-bit quantisation with quality and context trade-offs.

That closes the gap this slot flagged a month ago, when GLM-5.3 matched Kimi K3's score of 60 on the Artificial Analysis Intelligence Index but had no released weights. The "weights actually released" criterion is now satisfied. What has not changed is the quality picture: matching K3 is not beating it, and nothing in this cycle's evidence is an independent head-to-head evaluation putting GLM-5.3 ahead on the slot's stated criteria. Our bar for a title change is an independent result showing the challenger wins, so K3 keeps the slot.

Two cautions for anyone planning around this. The Hugging Face model card tags GLM-5.3 as license:other, not MIT — the MIT terms noted previously applied to the smaller GLM-5.3-Flash checkpoint, and the flagship's licence should be read directly before commercial use. And the local-hardware figures above come from a single social-media summary, not a measured serving benchmark, so treat them as an order-of-magnitude guide rather than a costed comparison against K3's 2.8T-parameter, 104B-active footprint.

We will revisit as soon as a third-party board publishes a separated score, or a like-for-like throughput and cost comparison, for the released GLM-5.3 checkpoint against K3.

evidence: [1] [2] [3]
27 Aug 2026
challenged

Fal's post-trained MiniMax H3 Max enters AA's top three, two points off the pick

Artificial Analysis's with-audio text-to-video board has tightened again. As of 27 August the top five reads Alibaba Wan 3.0 on 1,240, Gemini Omni Flash on 1,237, Fal's post-trained MiniMax H3 Max on 1,235, MiniMax H3 Open Weights on 1,227 and Dreamina Seedance 2.0 720p on 1,221. A week earlier the same board had Wan 3.0 on 1,247 and the pick on 1,239, so the nominal gap at the top has narrowed from eight points to three.

The new name is Fal MiniMax H3 Max, a post-train of MiniMax's model by fal, arriving straight into third. Three models now sit inside five Elo points of each other, which on this board is a tie rather than a ranking. There is no independent evidence in this cycle on the other slot criteria — clip length, character consistency, image-to-video — for any of the three, and the pick retains its lead on the no-audio board and on arena.ai.

Google also published Gemini Omni 1.1 Flash, framed around more build-time control. That is a vendor post with no third-party scores attached, so it does not shift the pick's standing or its recorded version; we will wait for arena and AA placements before treating 1.1 as the tracked build.

No change to the title. Wan 3.0 and MiniMax H3 remain the standing challengers, now joined by fal's post-train, and the audio-track weakness flagged in the caveats is unresolved.

evidence: [1] [2]
27 Aug 2026
challenged

Qwen3.8-Flash-Next arrives with day-two GGUFs but no independent scores yet

A new Qwen entry, Qwen3.8-Flash-Next, went up on Hugging Face on 24 August and is trending, with an unsloth GGUF conversion following on 26 August (model, GGUFs). Both cards list it as an image-text-to-text model, so it matches Gemma 4 26B A4B's multimodal input, and the existence of community quantisations within two days means the practical questions for this slot — file size at 4-bit, headroom for KV cache, whether function calling survives quantisation — are at least testable now rather than hypothetical.

That is as far as it goes. There is no independent evaluation of Flash-Next in evidence: it does not appear on Artificial Analysis's small open-source board as of the 22 August snapshot, which still shows Qwen3.8 27B at 52/44/43 across reasoning efforts, Qwen3.6 27B (Reasoning) at 38, Muse Glimmer at 35 and this pick at 26 (AA). No file sizes are published in the material we have, so the 16GB fit is unverified, and the Hugging Face card records the licence only as "other" — a step down in clarity from the incumbent's Apache-2.0.

The incumbent keeps the title on the same grounds as before: official Q4_0 QAT builds at 14.6GB that actually load on a 16GB machine, with native tool calling. Gemma 4 26B A4B remains the weakest of the field on raw index score, and this slot is under active review — but a trending upload with no third-party numbers does not move it.

evidence: [1] [2] [3]
27 Aug 2026
challenged

GLM-5.3-Flash lands under MIT, but no independent scores yet

Two smaller open-weight releases turned up on Hugging Face this week: Z.ai's GLM-5.3-Flash, tagged MIT, and Alibaba's Qwen3.8-Flash-Next, tagged license:other. Both picked up community GGUF conversions from Unsloth within a day or two (GLM, Qwen), which is the usual signal that people are actually running them locally.

Neither displaces Kimi K3. These are Flash-class checkpoints, and the dossier contains no independent evaluation placing either above K3 on the criteria this slot is judged against. What is notable is the licence: GLM-5.3-Flash ships as MIT, against K3's custom Moonshot terms with their MAU and revenue thresholds. If the full GLM-5.3 weights follow on the same licence — the release was previously expected around mid-August — that becomes a serious challenge on licence and serving cost simultaneously, assuming the quality holds. On the evidence here it is a Flash variant only, so it goes on the board as a challenger rather than a contender.

Meanwhile the incumbent's position on serving cost has quietly improved. vLLM v0.28.0 headlines a Kimi-K3 optimisation push including Decode Context Parallel support and fused FlashKDA decode and prefill kernels. That does not fix K3 being enormous to host, but it narrows the practical gap that made the pick uncomfortable.

No change. Reviewing again when a third-party score for any GLM-5.3 checkpoint appears.

evidence: [1] [2] [3] [4] [5]
26 Aug 2026
challenged

MAI-Image-2.6 (preview) closes to 20 points of GPT Image 2 on Artificial Analysis

Artificial Analysis's text-to-image arena has re-ordered its top five, and the margin behind GPT Image 2 has tightened. The board now reads: GPT Image 2 (high) first at 1371, Microsoft AI's MAI-Image-2.6-Preview second at 1351, Reve 2.1 third at 1322, Google's Nano Banana 2 (Gemini 3.1 Flash Image Preview) fourth at 1321, and GPT Image 1.5 (high) fifth at 1310.

That puts 20 points between the pick and its nearest rival, down from the roughly 48-point cushion recorded at the last review, when Reve 2.1 was the closest challenger on this board. MAI-Image-2.6 has also moved past Reve into second on Artificial Analysis, matching the position it already held on Arena's text-to-image board.

No change to the title. The pick still leads the board, and nothing in this evidence shows the challenger ahead on photoreal quality, editability, text rendering or instruction following — the leaderboard movement is a shrinking deficit, not a lead. MAI-Image-2.6 also remains a preview release, so its scores and availability should be treated as provisional.

The existing caveat still stands: editing is where GPT Image 2 is weakest relative to the field, and the arrival of a stronger Microsoft entrant on the text-to-image side makes it worth watching whether the same model displaces MAI-Image-2.5-Pro at the top of the editing arena. If the text-to-image gap closes further on a settled, non-preview release, this slot is live.

evidence: [1]
24 Aug 2026
held

Weekly sweep: 10 slots reviewed, no changes

58 candidate event(s) examined across all slots; none met promotion criteria. All last-reviewed dates refreshed.

23 Aug 2026
challenged

Orze-ASR-3Way takes #1 on Open ASR; ARK-ASR-3B holds the title pending licence checks

The Open ASR leaderboard has a new leader: bosonai/Orze-ASR-3Way at 3.81 mean WER, ahead of ARK-ASR-3B on 4.76, MOSS-Transcribe-preview-2B on 4.87, MOSS-Transcribe-Diarize on 5.17 and Cohere Transcribe on 5.42 (Open ASR). That is close to a full WER point over the current pick, measured by an independent scorer rather than a vendor card, so it is a serious challenge on the slot's first criterion.

It is not yet enough to take the title. This slot judges licence and language coverage alongside WER, and we have no confirmation of Orze's licence terms, weight release or language list, nor any throughput figure. ARK's case here was never a single number — it was Apache-2.0 plus 19 languages plus competitive accuracy — and swapping it out on a leaderboard row alone would be the sort of move this tracker exists to avoid. We will revisit once the model card and an independent throughput run are available.

One useful correction in the same update: ARK-ASR-3B now appears on the leaderboard in its own right at 4.76, which retires the standing caveat that its numbers were card-derived rather than board-verified. The figure sits between the card's claimed 5.04% and AutoArk's rerun at 5.13%, and slightly ahead of both.

The Artificial Analysis speech-to-text movement this week — Fun-Realtime-ASR-preview at 1.7%, Scribe v2, the Azure MAI-Transcribe pair and Smallest AI Pulse Pro (AA) — is all commercial API territory and does not bear on an open-weight slot.

evidence: [1] [2] [3]
23 Aug 2026
challenged

Claude Opus 5 tops Vellum's leaderboard, tightening the pressure on Fable 5

Vellum's public LLM leaderboard, refreshed on 16 August, now puts Claude Opus 5 in first place with a composite score of 64.7, ahead of Claude Mythos 5 (64.5), Claude Opus 4.8 (57.9), Claude Sonnet 5 (57.4) and Kimi K3 (56.0). Claude Fable 5, the current title holder, does not appear in that top five.

That is a genuine data point for Opus 5, and it lines up with the picture we already recorded: Artificial Analysis scores the two as effectively tied (63 against 62) at half the list price, and Opus 5 sits inside the chasing cluster on Arena preference voting.

It is not yet enough to move the title. Fable 5's absence from Vellum's top five is ambiguous — the board does not tell us whether it was evaluated and fell short or simply is not covered — and a single composite leaderboard cannot outweigh Fable's continued first place on Arena text and on the official Terminal-Bench 2.1 board (83.8% with Claude Code at xhigh effort). The slot's criteria weight long-context reliability and multimodal capability too, neither of which this evidence speaks to.

The practical reading for anyone choosing today is unchanged but sharpening: Fable 5 remains first-among-equals at the top of a statistical tie, while Opus 5 is now the pick that at least one independent board ranks above every other Anthropic model. If a second independent evaluation shows Opus 5 clearly ahead of Fable on reasoning or long-context work, this title changes hands.

evidence: [1]
23 Aug 2026
held

Weekly review: all 9 titles re-verified

All 9 picks were re-verified against the live leaderboards, repositories and release pages. 19 details moved with the evidence and the affected pick pages have been updated.

22 Aug 2026
challenged

Qwen3.8 27B extends its lead on AA's small board but still has no verified 16GB build

Artificial Analysis's small open-source board now has Qwen3.8 27B occupying the top three places, with its highest reasoning-effort setting scoring 52, followed by medium at 44 and low at 43. Qwen3.6 27B (Reasoning) drops to #4 on 38 and Muse Glimmer (high) to #5 on 35. The current pick, Gemma 4 26B A4B, sits well down the board at 26.

On raw quality this is no longer close. But this slot is not scored on quality alone: it asks for quality per GB, tool-calling that survives quantisation, and a model that actually runs in 16GB of RAM. On those points nothing has changed. There is still no official quantised release for Qwen3.8 27B with published file sizes, and no independent test of its function calling at 4-bit. The precedent is not encouraging — the dense Qwen3.6 27B's 4-bit GGUFs land between 15.4GB and 16.8GB, which leaves nothing for KV cache on a 16GB machine, and Qwen3.8 27B is the same size class.

The one new artefact in this cycle is z-lab/Qwen3.8-27B-DFlash2, a block-diffusion draft model for speculative decoding, tagged for sglang and vLLM. That is a throughput trick for server deployments, not evidence of a local 16GB path.

Gemma 4 26B A4B holds the title on its 14.6GB official Q4_0 QAT build. If Qwen ships an official quantised release that fits with room for context, this changes quickly.

evidence: [1] [2]
20 Aug 2026
challenged

Alibaba Wan 3.0 takes #1 on Artificial Analysis text-to-video, eight points ahead of the pick

Artificial Analysis's text-to-video leaderboard has a new leader: Alibaba Wan 3.0 enters at #1 with an Elo of 1,247, pushing Gemini Omni Flash to second on 1,239. The rest of the top five is MiniMax H3 Open Weights (1,228), Dreamina Seedance 2.0 720p (1,221) and Wan2.7-260612 (1,156).

That is an independent evaluation and a real lead change, so Wan 3.0 goes on the board as a challenger. It is not yet enough to take the title. The margin is eight Elo points on a crowd-vote board where the pick's own score has drifted by a couple of points since we last checked — that is inside the noise you would expect from arena scoring, and it says nothing about the three other criteria this slot judges on. We have no independent figures for Wan 3.0 on native audio, maximum clip length, or character consistency across shots, and no image-to-video placement, where Gemini Omni Flash still leads the no-audio board at 1,368 and sits close behind Seedance 2.0 on the with-audio board.

The practical read: if your work is single-shot text-to-video and you can run your own comparison, Wan 3.0 is now worth testing first. If you depend on synced audio in the same pass or on multi-turn editing that holds a character across changes, there is no evidence yet that Wan 3.0 matches the incumbent. We will revisit once AA's image-to-video and with-audio boards pick it up, or once a second independent source reports on its audio and clip length.

evidence: [1]
20 Aug 2026
held

Wan 3.0 takes the with-audio video crown; Qwen3.8-27B posts its first real score

Today's spot-check against the live boards moved several numbers. The biggest: Alibaba's Wan 3.0 now leads Artificial Analysis's with-audio text-to-video table at 1,247, with Gemini Omni Flash second on 1,239 — though the pick still tops the silent board at 1,323 and arena.ai's Text-to-Video Arena at 1,512, so the title holds, under real pressure. MiniMax H3, meanwhile, leads arena.ai's Image-to-Video Arena outright at 1,489.

Qwen3.8-27B has its first independent score, and it is a statement: 52 on Artificial Analysis's small open-source board — top of the table, double the current on-device titleholder's 26. It graduates from unverified to early data on the radar, and the on-device title is now under open review.

At the frontier, Artificial Analysis has nudged both leaders up: Claude Opus 5 (max) to 63 and Claude Fable 5 to 62 — still effectively tied, still a two-horse Claude race at the top. And one provenance correction in open speech-to-text: ARK-ASR's board-topping claim traces to its model card, not the leaderboard's public export, and its page now says so plainly.

evidence: [1] [2] [3] [4]
19 Aug 2026
challenged

Qwen3.8 27B quantised builds appear, but the 16GB question is still open

Artificial Analysis's small open-source board still puts Qwen3.8 27B first with a score of 52, well ahead of the current pick's 26, with Qwen3.6 27B (Reasoning) at 38 and Muse Glimmer (high) at 35. That ranking has not moved this week, and on raw quality the pick is clearly outclassed.

What is new is supply rather than evidence: a community GGUF conversion of Qwen3.8 27B is trending on Hugging Face, tagged for llama.cpp, tool-calling, long context and multimodal input. Two caveats keep it from shifting the title. First, it is an "abliterated" derivative — a safety-modified community remix, not a candidate we would seat as a title holder. Second, the listing carries no per-quantisation file sizes and no evaluation, so it says nothing about whether a 27B dense model at 4-bit leaves room for KV cache on a 16GB machine, or whether its function calling survives quantisation. Those are two of the three criteria for this slot.

So the position is unchanged: Gemma 4 26B A4B keeps the title on deployability, with official Q4_0 QAT builds around 14.6GB and native tool calling, while Qwen3.8 27B holds the quality crown and an unproven fit. The moment someone publishes independent quantised sizing plus tool-calling results for a first-party Qwen3.8 27B GGUF at 16GB, this slot is likely to turn over.

evidence: [1] [2]
17 Aug 2026
held

Weekly sweep: 10 slots reviewed, no changes

60 candidate event(s) examined across all slots; none met promotion criteria. All last-reviewed dates refreshed.

16 Aug 2026
challenged

Claude Fable 5 tops official Terminal-Bench 2.1 board; Opus 5 holds Arena WebDev #1

Two independent boards moved today, and they point in different directions.

On the Arena WebDev leaderboard, Anthropic's claude-opus-5-max is the new #1 at 1692, with kimi-k3-max second at 1674, qwen3.8-max third at 1667, claude-opus-5-high fourth at 1663 and grok-4.6-high fifth at 1631. That keeps the incumbent in front on human-preference web coding, and leaves the Kimi K3 challenger note intact: still the strongest open-weight showing, roughly 18 points off the top.

The more interesting item is the official Terminal-Bench 2.1 board, where Fable 5 has taken #1 at 83.8%, ahead of GPT-5.5 at 83.1%, a second Fable 5 configuration at 80.4%, Grok 4.5 at 79.3% and Opus 4.8 at 78.9%. Opus 5 does not appear in that top five at all. We are treating this as a challenge rather than a title change for two reasons. First, these figures come from a different harness to the vals.ai Terminal-Bench 2.1 run cited on the card, so the numbers are not interchangeable and there is no clean head-to-head between Fable 5 and Opus 5 here. Second, the remainder of the top five is made up of prior-generation models, which suggests the board's coverage of the newest releases is still partial rather than that Opus 5 has been beaten on merit.

Practical read: Opus 5 stays the default for repo-scale agentic work, but terminal-heavy workloads remain the contested area, and Fable 5 now has an independent board lead there — at roughly double the token price. Nothing in today's evidence touches the SWE-bench Verified or tool-calling picture.

evidence: [1] [2]
16 Aug 2026
held

Weekly review: all 10 titles re-verified

All 10 picks were re-verified against the live leaderboards, repositories and release pages. 22 details moved with the evidence and the affected pick pages have been updated.

15 Aug 2026
held

All ten titles hold — and the frontier tightens to a three-way tie

No title changed hands today, but the gaps closed. At the frontier, Claude Fable 5, its cheaper sibling Claude Opus 5 and xAI's new Grok 4.6 now sit within about a point of each other on the independent indexes — with Grok at a fraction of the price. Gemini Omni Flash kept text-to-video but has slipped to third on image-to-video, behind Seedance 2.0 and MiniMax H3. And the small-model title is under real pressure: Qwen3.6 27B and Meta's Muse Glimmer both out-score Gemma 4 26B A4B, while squeezing the 16GB limit in different ways. Grok 4.6 and NVIDIA's Nemotron 3.5 Lightning join the radar.

evidence: [1] [2]
15 Aug 2026
added

NVIDIA's unnoticed 550B teachers, Writer's GLM remix, and three more join the radar

Five arrivals worth watching. The quietest is the biggest: NVIDIA published five domain-specialised Nemotron Labs Teacher models — 550B total, 55B active, permissively licensed — claiming parity with DeepSeek V4 Pro, and almost nobody noticed; downloads are still in the hundreds. Writer's Palmyra X6 is notable for the opposite reason: it is a post-trained remix of Z.ai's open GLM-5.2, which makes an open Chinese base model the substrate of a Western enterprise flagship.

The rest: FireRedTTS3 brings Apache-2.0 voice cloning across 24 languages straight at the open TTS title; ByteDance's Seed-2.0-Code pitches at the coding harnesses for a fraction of frontier prices; and Liquid AI's LFM2.5-VL-3B squeezes vision and tool use into 3.3GB on a laptop.

One correction from the same review: GLM-5.3's entry previously said its weights were held back for safety hardening. We can find no source for that — Z.ai has announced no weights release at all, and the model remains subscription-only. Its entry now says so.

evidence: [1] [2] [3] [4] [5]
15 Aug 2026
held

Sol takes Terminal-Bench, DeepSeek closes on coding, Reve edges image editing

Real movement beneath the titles this week. On agentic terminal work, GPT-5.6 Sol now leads the independent Terminal-Bench 2.1 runs at 85.8%, ahead of Claude Opus 5 at 84.6% — and DeepSeek V4 Pro has closed to 96.4% on SWE-bench Verified against Opus 5's 97.0%, as an open-weight model at a fraction of the cost per task. In image editing, Reve 2.1 has edged ahead of GPT Image 2 on the Artificial Analysis arena — a statistical tie, and GPT Image 2 still holds both Arena boards. Fresh 70B head-to-heads put SGLang slightly ahead of vLLM at every tested concurrency, tightening that race too.

Two clarifications for the record: Qwen3.8's open checkpoint is text-only with a 262K native context — vision and the 1M window stay on the hosted Max — and Google's gemini-embedding-2 has been generally available since April. Every affected pick's page has been updated.

evidence: [1] [2] [3]
15 Aug 2026
challenged

Muse Spark closes on Fable, Kimi K3 joins the coding chase

The chasing pack behind four titles has reshuffled. Meta's Muse Spark 1.2 now sits fourth on the Arena text board at 1,498 — the nearest non-Claude rival to Fable 5's 1,506, and closer than Grok 4.6 gets on any board, so it takes Grok's place among the overall challengers. On coding, Kimi K3 has climbed to second on the WebDev arena at 1,674, above GPT-5.6 Sol's 1,622 — an open-weight model within twenty points of Opus 5 — and joins the challenger list there.

In video, Black Forest Labs' FLUX-3 Video has climbed to second on the Arena text-to-video board at 1,496±17, inside Gemini Omni Flash's error band, and joins MiniMax H3 in that chase. Among small models, Muse Glimmer 30B takes over as lead challenger to Gemma 4 — its independent score already sits above the titleholder's, though its widest claimed margins are vendor-run against Gemma's heavier dense sibling. And in open speech-to-text, MOSS-Transcribe-Diarize — the production sibling of the preview model already chasing ARK-ASR — joins the list with built-in diarisation and the adoption its sibling never found.

Every title stays where it was. The overall crown is the closest race: Claude Opus 5 now leads on the Artificial Analysis index and Humanity's Last Exam while Fable 5 keeps the human-preference crown by a nose — a genuine split decision worth watching.

evidence: [1] [2] [3] [4]
14 Aug 2026
added

The board, corrected: eight of ten titles change

Claude Fable 5 takes best overall LLM — Arena #1 since 1 July, though the top is genuinely tight and its own cheaper sibling runs it close on the Artificial Analysis index. Kimi K3 takes best open-weight LLM with the highest score an open model has ever posted there. Claude Opus 5 holds best for coding from its 24 July launch, when it swept SWE-bench Verified and the WebDev arena. Gemma 4 26B A4B is the small-model pick — a sparse MoE that fits 16GB of RAM without giving up tool calling. In speech, Fish Audio S2 Pro has led open TTS since March (mind the non-commercial licence) and ARK-ASR-3B tops the Open ASR Leaderboard at 4.76% WER. GPT Image 2 has owned both image arenas since April; Gemini Omni Flash holds video by a nose over MiniMax H3; Nemotron-3-Embed-8B leads retrieval embeddings on RTEB; and vLLM keeps inference serving, dated from its real June 2023 debut.

evidence: [1] [2] [3] [4] [5] [6]
14 Aug 2026
added

RightSignal goes live: ten titles, one working answer each

The board opens with ten titles across language models, speech, image & video, and retrieval & tooling. The idea is simple: one current, defensible answer per category — the model we'd actually reach for today — with the reasoning on show. Every change from here on lands as a dated entry in this changelog, with the evidence beside it.

27 Jul 2026
changed

Kimi K3 takes best open-weight LLM

The world's first open 3T-class model (2.8T total, 104B active), trading blows with Claude Fable 5 and GPT-5.6 Sol rather than with other open weights.

Official release: Kimi K3

Recorded retrospectively in the August 2026 backfill.

evidence: [1]
24 Jul 2026
changed

Claude Opus 5 takes best LLM for coding

New state of the art on SWE-bench Verified and the WebDev arena at half Fable 5's price — which made it the practical default for long-running agents on launch day.

Official release: Claude Opus 5

Recorded retrospectively in the August 2026 backfill.

evidence: [1]
16 Jul 2026
changed

Nemotron-3-Embed-8B takes best text embeddings

Number one on multilingual RTEB — the contamination-resistant board with private held-out sets — at 78.46 average nDCG@10 across 16 benchmarks. And unlike every proprietary model it displaced, the weights are downloadable under a permissive licence.

Recorded retrospectively in the August 2026 backfill.

evidence: [1] [2]
1 Jul 2026
changed

Claude Fable 5 takes best overall LLM

Number one on the Arena text leaderboard at 1506 Elo, having redeployed globally on 1 July after its 9 June launch was pulled back.

Official release: Claude Fable 5

Recorded retrospectively in the August 2026 backfill.

evidence: [1]
22 Jun 2026
changed

ARK-ASR-3B takes best open-weight STT

Took the Open ASR Leaderboard top spot at 4.76% WER under Apache-2.0 — the first open model clearly inside the range of the commercial transcription APIs.

Official release: ARK-ASR-3B

Recorded retrospectively in the August 2026 backfill.

evidence: [1]
16 Jun 2026
changed

GLM-5.2 takes best open-weight LLM

A 753B MIT-licensed release that retook the lead from DeepSeek-V4-Pro — the model Kimi K3 later had to benchmark itself against.

Official release: GLM-5.2

Recorded retrospectively in the August 2026 backfill.

evidence: [1]
11 May 2026
milestone

vLLM tops the Artificial Analysis provider board

First of all 12 measured providers on the Qwen 3.5 397B release, with sub-second time-to-first-token on 10,000-token prompts. The open engine is now the fastest option, not just the cheap one.

Recorded retrospectively in the August 2026 backfill.

evidence: [1]
21 Apr 2026
changed

GPT Image 2 takes best image generation

Putting a reasoning model in front of generation is what finally cracked complex multi-element prompts — it leads the image arena by a clear 44-point Elo margin.

Official release: GPT Image 2

Recorded retrospectively in the August 2026 backfill.

evidence: [1]
26 Feb 2026
changed

Nano Banana 2 takes best image generation

Pro-tier instruction following and text rendering at Flash latency and price, rolled straight into the Gemini app, Search AI Mode and Lens.

Official release: Nano Banana 2

Recorded retrospectively in the August 2026 backfill.

evidence: [1]
12 Feb 2026
changed

Seedance 2.0 takes best video generation

Fifteen-second clips with genuine camera control and a realism jump big enough that the Motion Picture Association and Disney went after it within days of release.

Official release: Seedance 2.0

Recorded retrospectively in the August 2026 backfill.

evidence: [1]
11 Feb 2026
changed

GLM-5 takes best open-weight LLM

A 744B/40B-active model claiming best-in-class among all open models on reasoning, coding and agentic work.

Official release: GLM-5

Recorded retrospectively in the August 2026 backfill.

evidence: [1]
5 Feb 2026
changed

Claude Opus 4.6 takes best overall LLM

Took the top spot back from Gemini 3 Pro by leading Humanity's Last Exam and fixing context rot — and stayed Anthropic's arena leader even after Opus 4.7 and 4.8 shipped.

Official release: Claude Opus 4.6

Recorded retrospectively in the August 2026 backfill.

evidence: [1]
5 Feb 2026
changed

Claude Opus 4.6 takes best LLM for coding

Pushed SWE-bench Verified to 81.42% and took the top Terminal-Bench 2.0 score; the 4.7 and 4.8 refreshes extended the same era rather than ending it.

Official release: Claude Opus 4.6

Recorded retrospectively in the August 2026 backfill.

evidence: [1]
15 Jan 2026
changed

voyage-4-large takes best text embeddings

The first production-grade MoE embedding model: 8.20% better general retrieval than both gemini-embedding-001 and Cohere Embed v4, 14.05% over OpenAI v3-large, at 40% lower serving cost than comparable dense models.

Recorded retrospectively in the August 2026 backfill.

evidence: [1]
16 Dec 2025
changed

GPT Image 1.5 takes best image generation

Fixed the premature cropping and warm colour cast that made GPT Image 1 output obvious, and ran roughly four times faster at 20% lower API cost.

Official release: GPT Image 1.5

Recorded retrospectively in the August 2026 backfill.

evidence: [1]
18 Nov 2025
changed

Gemini 3 Pro takes best overall LLM

Launched claiming the LMArena top spot outright at 1501 Elo — the first model past 1500 — alongside 91.9% GPQA Diamond.

Official release: Gemini 3 Pro

Recorded retrospectively in the August 2026 backfill.

evidence: [1]
18 Nov 2025
milestone

vLLM adopts a two-week release train

Release week opens every other Monday on the healthiest recent commit, with cherry-picks gated on full CI, performance benchmarks and accuracy evals — predictability over heroics.

Recorded retrospectively in the August 2026 backfill.

evidence: [1]
12 Nov 2025
changed

Step-Audio-EditX takes best open-weight TTS

Apache-2.0 weights that added iterative post-hoc editing of emotion and style on top of strong zero-shot cloning — the workflow feature practitioners had been faking with regeneration loops.

Official release: Step-Audio-EditX

Recorded retrospectively in the August 2026 backfill.

evidence: [1]
30 Sep 2025
changed

Sora 2 takes best video generation

Best-in-class physics and character consistency wrapped in a consumer app that put video generation in front of millions of non-specialists for the first time.

Official release: Sora 2

Recorded retrospectively in the August 2026 backfill.

evidence: [1]
26 Aug 2025
changed

Nano Banana takes best image generation

Google shipped character consistency and multi-image fusion that actually survived repeated edits — the thing that had been blocking real production use.

Official release: Nano Banana

Recorded retrospectively in the August 2026 backfill.

evidence: [1]
7 Aug 2025
changed

GPT-5 takes best overall LLM

The unified reasoning-plus-router release retook the arena from Gemini 2.5 Pro and set state of the art across maths, coding and multimodal benchmarks simultaneously.

Official release: GPT-5

Recorded retrospectively in the August 2026 backfill.

evidence: [1]
5 Aug 2025
changed

gpt-oss-20b takes best small / on-device LLM

21B total with 3.6B active and native MXFP4, engineered by OpenAI to run in exactly 16GB — the first model where on-device was the design constraint rather than an afterthought.

Official release: gpt-oss-20b

Recorded retrospectively in the August 2026 backfill.

evidence: [1]
17 Jul 2025
changed

Canary Qwen 2.5B takes best open-weight STT

Bolting an LLM decoder onto the Canary encoder hit 5.63% WER and gave transcription plus summarisation in one pass; nothing open beat it for nine months.

Official release: Canary Qwen 2.5B

Recorded retrospectively in the August 2026 backfill.

evidence: [1]
14 Jul 2025
changed

gemini-embedding-001 takes best text embeddings at GA

Top of the MTEB Multilingual board continuously since its March experimental launch, and genuinely deployable from GA: 100+ languages, 3072 dimensions with Matryoshka truncation, on a production-grade API.

Recorded retrospectively in the August 2026 backfill.

evidence: [1]
11 Jul 2025
changed

Kimi K2 Instruct takes best open-weight LLM

A 1T-parameter MoE with 32B active claiming state of the art on agentic and coding benchmarks against Claude Opus 4 and GPT-4.1 — not merely against other open models.

Official release: Kimi K2 Instruct

Recorded retrospectively in the August 2026 backfill.

evidence: [1]
28 May 2025
changed

Chatterbox takes best open-weight TTS

The first open model with credible blind-test evidence of beating ElevenLabs (63.75% preference), shipped MIT with emotion-exaggeration control.

Official release: Chatterbox

Recorded retrospectively in the August 2026 backfill.

evidence: [1]
22 May 2025
changed

Claude Opus 4 takes best LLM for coding

Billed as the world's best coding model on 72.5% SWE-bench Verified, and held the practitioner crown through GPT-5's August challenge — the benchmark gap stayed inside noise while agentic harnesses stayed Claude-default.

Official release: Claude Opus 4

Recorded retrospectively in the August 2026 backfill.

evidence: [1]
20 May 2025
changed

Veo 3 takes best video generation

Native synchronised audio with dialogue and lip sync ended the silent-film era of AI video in one release — no competitor had an answer for months.

Official release: Veo 3

Recorded retrospectively in the August 2026 backfill.

evidence: [1]
6 May 2025
milestone

vLLM joins the PyTorch Foundation

Governance moves out of UC Berkeley onto neutral footing at 46,500 stars and 1,000+ contributors — one of the first projects under the newly expanded umbrella foundation.

Recorded retrospectively in the August 2026 backfill.

evidence: [1]
29 Apr 2025
changed

Qwen3-14B takes best small / on-device LLM

Added switchable thinking mode at a size that still quantises into 16GB, and swept the small-model benchmarks at launch.

Official release: Qwen3-14B

Recorded retrospectively in the August 2026 backfill.

evidence: [1]
15 Apr 2025
changed

Kling 2.0 takes best video generation

Briefly the strongest model you could actually buy access to worldwide, while Veo 2 was still gated behind waitlists and regional limits.

Official release: Kling 2.0

Recorded retrospectively in the August 2026 backfill.

evidence: [1]
25 Mar 2025
changed

Gemini 2.5 Pro takes best overall LLM

Debuted directly at #1 on LMArena and turned the experimental streak into a stable, generally available product with a 1M-token context and built-in thinking.

Official release: Gemini 2.5 Pro

Recorded retrospectively in the August 2026 backfill.

evidence: [1]
12 Mar 2025
changed

Gemma 3 12B takes best small / on-device LLM

Brought 128k context and vision to the 16GB tier, with official QAT int4 checkpoints — the quantised build was the intended artefact, not a community guess.

Official release: Gemma 3 12B

Recorded retrospectively in the August 2026 backfill.

evidence: [1]
24 Feb 2025
changed

Claude 3.7 Sonnet takes best LLM for coding

The first hybrid reasoning model, hitting 70.3% on SWE-bench Verified — and it shipped Claude Code alongside, which moved the goalposts from snippets to repo-scale work.

Official release: Claude 3.7 Sonnet

Recorded retrospectively in the August 2026 backfill.

evidence: [1]
27 Jan 2025
milestone

vLLM's V1 engine ships in alpha

A re-architected core behind a single environment variable: up to 1.7x higher throughput purely from CPU overhead reduction, with the GPU kernels essentially unchanged.

Recorded retrospectively in the August 2026 backfill.

evidence: [1]
20 Jan 2025
changed

DeepSeek-R1 takes best open-weight LLM

Put open weights on the reasoning curve for the first time, under MIT, and held the crown through the whole first half of 2025.

Official release: DeepSeek-R1

Recorded retrospectively in the August 2026 backfill.

evidence: [1]
8 Jan 2025
changed

Phi-4 takes best small / on-device LLM

Synthetic-data training gave this 14B maths and reasoning scores well above its weight class, with MIT weights.

Official release: Phi-4

Recorded retrospectively in the August 2026 backfill.

evidence: [1]
7 Jan 2025
changed

voyage-3-large stretches Voyage's lead

9.74% ahead of OpenAI's v3-large across 100 datasets — and its 512-dimension binary vectors beat OpenAI's 3072-dimension floats outright at roughly 200x less storage. The cost argument stopped being close.

Recorded retrospectively in the August 2026 backfill.

evidence: [1]
26 Dec 2024
changed

Kokoro-82M takes best open-weight TTS

An 82M-parameter model that sounded better than things twenty times its size and ran on CPU — roughly 12M downloads and the default for anyone shipping TTS on a budget.

Official release: Kokoro-82M

Recorded retrospectively in the August 2026 backfill.

evidence: [1]
26 Dec 2024
changed

DeepSeek-V3 takes best open-weight LLM

A 671B MoE that beat the 405B at a fraction of the serving cost, and flipped local-LLM consensus to the Chinese labs essentially overnight.

Official release: DeepSeek-V3

Recorded retrospectively in the August 2026 backfill.

evidence: [1]
16 Dec 2024
changed

Veo 2 takes best video generation

4K output and a markedly better grasp of physics and cinematographic language, winning head-to-head comparisons against every rival at launch.

Official release: Veo 2

Recorded retrospectively in the August 2026 backfill.

evidence: [1]
6 Dec 2024
changed

Gemini-Exp-1206 takes best overall LLM

Google's experimental checkpoint line took the top spot off GPT-4o and never handed it back through Q1 2025, despite brief Grok 3 and GPT-4.5 challenges.

Official release: Gemini-Exp-1206

Recorded retrospectively in the August 2026 backfill.

evidence: [1]
4 Dec 2024
challenged

SGLang v0.4 emerges as the serious challenger

A zero-overhead batch scheduler, cache-aware load balancer, and data-parallel attention for DeepSeek models — the point where vLLM's title stopped being uncontested.

Recorded retrospectively in the August 2026 backfill.

evidence: [1]
30 Oct 2024
changed

Recraft V3 takes best image generation

The first time the top spot was settled by a public blind-vote board rather than vibes — #1 on the Artificial Analysis image arena at 1172 Elo, ahead of Midjourney and OpenAI.

Official release: Recraft V3

Recorded retrospectively in the August 2026 backfill.

evidence: [1]
7 Oct 2024
changed

F5-TTS takes best open-weight TTS

Flow matching gave noticeably cleaner prosody and faster inference than XTTS; the community moved on quickly despite the non-commercial licence.

Official release: F5-TTS

Recorded retrospectively in the August 2026 backfill.

evidence: [1]
18 Sep 2024
changed

voyage-3 takes best text embeddings

Beat text-embedding-3-large by 7.55% on retrieval across domains at 2.2x lower price, with 1024 dimensions instead of 3072 and a 32k context. The nominal MTEB #1 at the time, NV-Embed-v2, was CC-BY-NC and therefore off the table commercially.

Recorded retrospectively in the August 2026 backfill.

evidence: [1]
5 Sep 2024
milestone

vLLM v0.6.0: the performance overhaul

The CPU-side bottlenecks get attacked — API server split from the engine over ZMQ, multi-step scheduling, async output processing — for 2.7x throughput and 5x faster time-per-output-token on Llama 3 8B.

Recorded retrospectively in the August 2026 backfill.

evidence: [1]
20 Jun 2024
changed

Claude 3.5 Sonnet takes best LLM for coding

Flipped practitioner consensus away from OpenAI overnight by solving 64% of agentic coding problems against Claude 3 Opus's 38%; the October refresh made it the default in Cursor and friends.

Official release: Claude 3.5 Sonnet

Recorded retrospectively in the August 2026 backfill.

evidence: [1]
17 Jun 2024
changed

Runway Gen-3 Alpha takes best video generation

A large jump in temporal consistency and motion fidelity that shipped to paying users — which mattered more than Sora's February demo reel nobody outside OpenAI could touch.

Official release: Runway Gen-3 Alpha

Recorded retrospectively in the August 2026 backfill.

evidence: [1]
13 May 2024
changed

GPT-4o takes best overall LLM

Reclaimed the crown for OpenAI at launch — after its anonymous "gpt2-chatbot" arena run broke records pre-announcement — and made native multimodality table stakes.

Official release: GPT-4o

Recorded retrospectively in the August 2026 backfill.

evidence: [1]
26 Mar 2024
changed

Claude 3 Opus takes best overall LLM

The first non-OpenAI model ever to take #1 on Chatbot Arena, breaking a year of GPT-4 dominance on stronger reasoning and genuine long-form fluency.

Official release: Claude 3 Opus

Recorded retrospectively in the August 2026 backfill.

evidence: [1]
7 Feb 2024
changed

Canary 1B takes best open-weight STT

Took the top of the Open ASR Leaderboard off Whisper at a third of the size — and NVIDIA held that spot for over a year.

Official release: Canary 1B

Recorded retrospectively in the August 2026 backfill.

evidence: [1]
25 Jan 2024
changed

text-embedding-3-large takes the title back for OpenAI

OpenAI's answer to the open leaders: a real MTEB jump over ada-002 plus Matryoshka dimension truncation, letting teams cut vector-database cost by shortening vectors instead of re-embedding — while keeping the zero-ops API path.

Recorded retrospectively in the August 2026 backfill.

evidence: [1]
21 Dec 2023
changed

Midjourney v6 takes best image generation

Trained from scratch over nine months, v6 closed the prompt-adherence gap while keeping the aesthetic lead; the v6.1 web release held the crown through a quiet year.

Official release: Midjourney v6

Recorded retrospectively in the August 2026 backfill.

evidence: [1]
31 Oct 2023
changed

XTTS-v2 takes best open-weight TTS

Made six-second zero-shot voice cloning across 17 languages actually work, and became the workhorse behind most open TTS products for the next year.

Official release: XTTS-v2

Recorded retrospectively in the August 2026 backfill.

evidence: [1]
1 Oct 2023
changed

DALL-E 3 takes best image generation

The first model that reliably drew what you actually asked for — and the ChatGPT integration meant iterating in plain language instead of fighting Discord parameter strings.

Official release: DALL-E 3

Recorded retrospectively in the August 2026 backfill.

evidence: [1]
12 Sep 2023
changed

bge-large-en-v1.5 takes best text embeddings

The first open model you could swap in for ada-002 and simply win: MTEB average 64.23 against 60.99, retrieval 54.29 against 49.25, at 335M parameters under MIT — one cheap GPU, no per-token bill. The era of the default OpenAI embedding ends here.

Recorded retrospectively in the August 2026 backfill.

evidence: [1]
20 Jun 2023
milestone

vLLM goes public with PagedAttention

Up to 24x the throughput of HuggingFace Transformers with no model changes — already battle-tested for two months behind Chatbot Arena and the Vicuna demo. The category's defining tool arrives.

Recorded retrospectively in the August 2026 backfill.

evidence: [1]