<?xml version="1.0" encoding="UTF-8"?>
<rss version="2.0" xmlns:atom="http://www.w3.org/2005/Atom">
<channel>
  <title>RightSignal — changelog</title>
  <link>https://rightsignal.co.uk/</link>
  <atom:link href="https://rightsignal.co.uk/feed.xml" rel="self" type="application/rss+xml"/>
  <description>Which AI model is actually best right now? One evidence-based pick per category, continuously reviewed — every change dated, cited and on the record.</description>
  <language>en-gb</language>
  <item>
    <title>[CHALLENGED] Wan 3.0&#39;s lead over Gemini Omni Flash widens to five Elo on AA&#39;s with-audio text-to-video board</title>
    <link>https://rightsignal.co.uk/slots/video-gen/</link>
    <guid isPermaLink="false">https://rightsignal.co.uk/changelog/2026-09-13-video-gen</guid>
    <pubDate>Sun, 13 Sep 2026 06:00:00 +0000</pubDate>
    <description>&lt;p&gt;Artificial Analysis&#39;s with-audio text-to-video board, read today, puts Alibaba Wan 3.0 first on 1242 with Gemini Omni Flash second on 1237, then fal&#39;s post-trained Minimax H3 Max on 1231 and MiniMax H3 Open Weights on 1225 (&lt;a href=&#34;https://artificialanalysis.ai/video/leaderboard/text-to-video&#34;&gt;AA text-to-video&lt;/a&gt;). That is the third successive read with Wan 3.0 at or ahead of the pick — level at 1239, then 1240 to 1239, now five points clear — so this is no longer a one-off reading, and the direction of travel is against the pick.&lt;/p&gt;
&lt;p&gt;We are not moving the title yet, for two reasons. First, the lead is confined to one board from one provider: the pick still tops AA&#39;s no-audio text-to-video table, where MiniMax H3 is second on 1301, and leads on image-to-video without audio on 1365. Five Elo points on a single with-audio table is thin corroboration against that.&lt;/p&gt;
&lt;p&gt;Second, two of the four things this slot judges on — character consistency and clip length — still have no independent read for Wan 3.0 at all. We would be swapping a model whose limits we know (Google&#39;s own card lists complete consistency across edits as open; 720p standard output with upscaling for higher resolutions) for one whose limits are simply unmeasured.&lt;/p&gt;
&lt;p&gt;If Wan 3.0 holds this lead into the next read, or picks up a second board or an independent consistency or clip-length result, the title moves.&lt;/p&gt;</description>
  </item>
  <item>
    <title>[CHALLENGED] Breeze TTS 2 slips to #9 as StepAudio 2.5 TTS edges ahead</title>
    <link>https://rightsignal.co.uk/slots/tts-open-weight/</link>
    <guid isPermaLink="false">https://rightsignal.co.uk/changelog/2026-09-13-tts-open-weight</guid>
    <pubDate>Sun, 13 Sep 2026 06:00:00 +0000</pubDate>
    <description>&lt;p&gt;The pick has lost ground. On the Artificial Analysis provider-voice board (fetched 13 September) BreezeBlue Breeze TTS 2 Open Weights reads 1202 Elo at #9, down from the 1215 and 7th place that won it the title on 9 September. Sitting just above it at 1207 is StepFun StepAudio 2.5 TTS (Aug 2026), with the rest of the top ten — Cartesia Sonic 3.6 (1276), Inworld Realtime TTS-2 (1243), Speechify Simba 3.2 (1237), Qwen-Audio-3.0-TTS-Plus (1234), VUI Luna (1228), Gemini 3.1 Flash TTS, ElevenLabs v3 — either hosted-only or unconfirmed for downloadable weights.&lt;/p&gt;
&lt;p&gt;StepFun is the obvious challenger: the Step-Audio line already appears in this slot&#39;s caveats as the Apache-2.0-coded option self-hosters fall back on. But we have nothing in front of us confirming that checkpoints for the 2.5 TTS release have actually shipped, nor its licence terms, time-to-first-audio or VRAM footprint — all stated criteria here. A five-point Elo gap on one board at one refresh is also thin corroboration. That is a challenge, not a handover.&lt;/p&gt;
&lt;p&gt;Also newly ingested: tencent/AuK is trending on Hugging Face as a text-to-speech pipeline with zero-shot cloning, editing and separation tags, but carries no independent quality or latency measurement yet, so it changes nothing today.&lt;/p&gt;
&lt;p&gt;Breeze TTS 2 keeps the title on its verified package — sub-40ms first audio, 0.32 RTF on a warmed H100, ~7.7 GiB eager inference — with the standing caveat that the weights remain research/non-commercial. We will revisit as soon as StepAudio 2.5&#39;s weights and licence are pinned down or the board reading holds.&lt;/p&gt;</description>
  </item>
  <item>
    <title>[HELD] Weekly review: all 9 titles re-verified</title>
    <link>https://rightsignal.co.uk/changes/</link>
    <guid isPermaLink="false">https://rightsignal.co.uk/changelog/2026-09-13-held</guid>
    <pubDate>Sun, 13 Sep 2026 06:00:00 +0000</pubDate>
    <description>&lt;p&gt;All 9 picks were re-verified against the live leaderboards, repositories and release pages. 22 details moved with the evidence and the affected pick pages have been updated.&lt;/p&gt;
&lt;p&gt;Under heightened review after this pass: GPT Image 2 (best image generation), Gemini Omni Flash (best video generation). A title only changes through full adjudication.&lt;/p&gt;</description>
  </item>
  <item>
    <title>[CHALLENGED] Wan 3.0 extends its Artificial Analysis text-to-video lead over Gemini Omni Flash to five points</title>
    <link>https://rightsignal.co.uk/slots/video-gen/</link>
    <guid isPermaLink="false">https://rightsignal.co.uk/changelog/2026-09-11-video-gen</guid>
    <pubDate>Fri, 11 Sep 2026 06:00:00 +0000</pubDate>
    <description>&lt;p&gt;Artificial Analysis&#39;s text-to-video board on 11 September puts Alibaba Wan 3.0 first on 1242 Elo, with our pick Gemini Omni Flash second on 1237. Fal&#39;s post-trained Minimax H3 Max is third on 1232, MiniMax H3 Open Weights fourth on 1225 and Dreamina Seedance 2.0 720p fifth on 1220.&lt;/p&gt;
&lt;p&gt;The useful detail is the trend rather than the single number. At our first read the two models were level on 1239; at the last read Wan 3.0 was one point ahead on 1240; it is now five points clear. That is the same board moving the same way three times, which is more than noise, and it is why the challenger note has been sharpened.&lt;/p&gt;
&lt;p&gt;It is still not enough to move the title. This is one board measuring one of our four criteria. Gemini Omni Flash remains top of AA&#39;s no-audio text-to-video table and top on no-audio image-to-video, and there is still no independent read on Wan 3.0&#39;s character consistency or clip length — half our criteria are simply unmeasured for the challenger. A five-point Elo gap is also the sort of margin that sits inside a ranked range rather than establishing a clear win.&lt;/p&gt;
&lt;p&gt;So: the pick holds, but on a narrower base than when we awarded it. Audio-inclusive generation is the weak flank, and if Wan 3.0 holds or extends this gap at the next read, or a second source corroborates it on consistency or clip length, the title should change hands.&lt;/p&gt;</description>
  </item>
  <item>
    <title>[CHALLENGED] GPT-Image-2.5-Sunburst challenges GPT Image 2 for best image generation</title>
    <link>https://rightsignal.co.uk/slots/image-gen/</link>
    <guid isPermaLink="false">https://rightsignal.co.uk/changelog/2026-09-11-image-gen</guid>
    <pubDate>Fri, 11 Sep 2026 06:00:00 +0000</pubDate>
    <description>&lt;p&gt;GPT-Image-2.5-Sunburst would take this title on today&#39;s evidence, but the case has not yet cleared our bar for a change, so it is recorded as the challenger for now. The adjudicator&#39;s reasoning:&lt;/p&gt;
&lt;p&gt;GPT Image 2 has held this title since 21 April, but it is now behind on both boards we track for this slot.&lt;/p&gt;
&lt;p&gt;Arena had already moved: GPT-Image-2.5-Sunburst leads text-to-image on 1421 against the incumbent&#39;s 1381, and — more important for the editability criterion — tops single-image edit on 1520 against 1461, with GPT-Image-2.5-Flare second on 1491. Today&#39;s Artificial Analysis fetch corroborates it. Its text-to-image arena now reads Flare (max) 1187, Sunburst (max) 1180, GPT Image 2 (high) 1171, dropping the incumbent to third on the board where it previously sat first. MAI-Image-2.6 is fourth on 1145 and Reve 2.1 fifth on 1127.&lt;/p&gt;
&lt;p&gt;Two challengers split the boards. Flare is nominally ahead of Sunburst on Artificial Analysis, but by only 7 Elo, while Sunburst leads Arena on both text-to-image and single-image edit. Sunburst is therefore ahead on more of this slot&#39;s criteria, including editing, and takes the title; Flare becomes the challenger and could well overtake it as the AA sample grows.&lt;/p&gt;
&lt;p&gt;Two caveats for buyers. Both 2.5 variants are recent arrivals on these boards, and the Flare/Sunburst margin on Artificial Analysis is well inside noise, so the ordering inside the 2.5 family is not settled. And cost remains a real consideration: Artificial Analysis prices GPT Image 2 at $211 per 1,000 images against $38.90 for MAI-Image-2.6, which still leads AA&#39;s image-editing arena on 1122. If you are editing at volume rather than chasing the top of the prompt-adherence charts, the cheaper Microsoft model remains the sane buy.&lt;/p&gt;</description>
  </item>
  <item>
    <title>[CHALLENGED] Wan 3.0 edges to nominal #1 on Artificial Analysis text-to-video, one Elo point clear</title>
    <link>https://rightsignal.co.uk/slots/video-gen/</link>
    <guid isPermaLink="false">https://rightsignal.co.uk/changelog/2026-09-10-video-gen</guid>
    <pubDate>Thu, 10 Sep 2026 06:00:00 +0000</pubDate>
    <description>&lt;p&gt;Artificial Analysis&#39;s text-to-video board now lists Alibaba Wan 3.0 at #1 on 1240 Elo, with Gemini Omni Flash second on 1239 — a one-point gap, where the two models were on an identical 1239 at the last read. Fal&#39;s post-trained Minimax H3 Max is third on 1235 and MiniMax H3 Open Weights fourth on 1228.&lt;/p&gt;
&lt;p&gt;One Elo point on a single board is not a lead in any practical sense; it is a tiebreak placing that could reverse on the next refresh. It is also not corroborated anywhere else in the dossier: no second independent board puts Wan 3.0 ahead, and nothing here reads on clip length or character consistency, which are two of this slot&#39;s four criteria. Native audio remains the softest part of the pick&#39;s case — on AA&#39;s with-audio tables the pick has been level with or behind Wan 3.0 on text-to-video and behind fal&#39;s post-trained Minimax H3 Max on image-to-video — but that pressure is unchanged rather than newly decisive.&lt;/p&gt;
&lt;p&gt;So the title holds and the challenger note is updated to record Wan 3.0&#39;s nominal top placing. Practical read for anyone choosing today: the top four on this board are close enough that prompt fit and workflow matter more than the ordering. If you need multi-turn conversational editing with audio in one pass, the pick is still the straightforward choice, with the caveat that Google&#39;s own known limitations list complete consistency across edits as unresolved. A second board putting Wan 3.0 clearly ahead, or a durable gap here, would take the title.&lt;/p&gt;</description>
  </item>
  <item>
    <title>[CHALLENGED] MiniCPM5-2B gets an official GGUF, tightening the challenge to Gemma 4 26B A4B</title>
    <link>https://rightsignal.co.uk/slots/llm-small/</link>
    <guid isPermaLink="false">https://rightsignal.co.uk/changelog/2026-09-10-llm-small</guid>
    <pubDate>Thu, 10 Sep 2026 06:00:00 +0000</pubDate>
    <description>&lt;p&gt;Gemma 4 26B A4B keeps the title, but one challenger has moved.&lt;/p&gt;
&lt;p&gt;OpenBMB has now published an official GGUF repo for MiniCPM5-2B, tagged for tool-calling, long context and on-device/edge use (&lt;a href=&#34;https://huggingface.co/openbmb/MiniCPM5-2B-GGUF&#34;&gt;repo&lt;/a&gt;). That answers the main objection in our previous note, which was that MiniCPM5-2B had no quantised build at all and so no on-device story. It is still not enough to take the slot: the dossier shows no published file sizes for those GGUFs, no independent 4-bit quality-retention numbers and no tool-calling-under-quantisation measurements. On Artificial Analysis&#39;s small open-models index MiniCPM5-2B scores 15, below this pick&#39;s 17, so the case for it rests entirely on quality per GB — credible for a sub-4B model, but not something an independent board has yet measured.&lt;/p&gt;
&lt;p&gt;The headline standings are unchanged. Qwen3.8 27B still leads aa-small-open at 34 (xhigh), with 28 medium, 26 low and 22 non-reasoning, against the pick&#39;s 17 at #10 (&lt;a href=&#34;https://artificialanalysis.ai/models/open-source/small&#34;&gt;standings&lt;/a&gt;). Qwen3.6 27B (Reasoning) has entered the top five at 22. But Qwen3.8 27B remains disqualified on this slot&#39;s memory criterion: the only official quantised release is the ~31GB FP8 GPU repo, and community 4-bit GGUFs of a 27B dense model land near 17GB, above a 16GB machine. Qwen3.8-Flash-Next&#39;s NVFP4 build is GPU-oriented with no listed size and an &#34;other&#34; licence.&lt;/p&gt;
&lt;p&gt;So the pick holds on deployability, not on raw score. A published, sized 4-bit build from either Qwen3.8 27B or MiniCPM5-2B, with tool-calling checked under quantisation, would likely settle this.&lt;/p&gt;</description>
  </item>
  <item>
    <title>[CHANGED] Claude Fable 5.1 replaces Claude Fable 5 as best overall LLM</title>
    <link>https://rightsignal.co.uk/slots/llm-overall/</link>
    <guid isPermaLink="false">https://rightsignal.co.uk/changelog/2026-09-10-llm-overall</guid>
    <pubDate>Thu, 10 Sep 2026 06:00:00 +0000</pubDate>
    <description>&lt;p&gt;The title moves within the family: Claude Fable 5.1 takes over from Claude Fable 5.&lt;/p&gt;
&lt;p&gt;Artificial Analysis re-based its Intelligence Index twice this month (v4.2 on 4 September, v4.3 on 7 September). On the current v4.3 scale, Claude Fable 5.1 at max and xhigh effort shares first place on 53 with GPT-6 Astra (max/xhigh), with Claude Opus 5 on 51 and Fable 5 itself down on 50. The official Terminal-Bench board has moved from 2.1 to v4.0, and there Fable 5.1 leads at 57.9%, with Opus 5 on 51.8% and Fable 5 third at 44.5% — the old 83.8% figure we were quoting describes a retired board. Vellum&#39;s composite of quality, speed and cost (6 September) also puts Fable 5.1 first at 65%.&lt;/p&gt;
&lt;p&gt;That leaves Arena text as the only board where Fable 5 still shows on top, at 1507 on the 2 September snapshot — but with a rank spread of 1-6 and Fable 5.1 (max) three points back at 1504, alongside Opus 4.6 (high) at 1505 and Opus 4.7 (high) at 1502. That is a tie band, not a lead worth holding a title on.&lt;/p&gt;
&lt;p&gt;GPT-6 Astra is the honest alternative and stays as challenger: it ties Fable 5.1 on the Index at less than half the cost per Index task ($3.26 against $7.63), and leads Artificial Analysis&#39;s own Terminal-Bench v4.0 runs at 59.1% to 52.0%. But it has no Arena text placement at all and sits seventh on Vellum at 57.2%, so Fable 5.1 is ahead on more of this slot&#39;s criteria. Anyone paying by the token should still price Astra.&lt;/p&gt;</description>
  </item>
  <item>
    <title>[CHANGED] GLM-5.3 replaces Kimi K3 as best open-weight LLM</title>
    <link>https://rightsignal.co.uk/slots/llm-open-weight/</link>
    <guid isPermaLink="false">https://rightsignal.co.uk/changelog/2026-09-10-llm-open-weight</guid>
    <pubDate>Thu, 10 Sep 2026 06:00:00 +0000</pubDate>
    <description>&lt;p&gt;Kimi K3 held this slot on a one-point lead over GLM-5.3 on Artificial Analysis&#39;s Intelligence Index v4.2. That scale has been retired. On v4.3 the lead is gone: AA&#39;s leaderboard fetched on 9 September has GLM-5.3 (max) at 45 and Kimi K3 (max) at 44, while AA&#39;s open-weights page on 7 September names the two as level at 44 at the top of the open field. On the most favourable reading for the incumbent, quality is now a tie.&lt;/p&gt;
&lt;p&gt;The slot&#39;s other criteria then decide it. The pick&#39;s own card already conceded that GLM-5.3 costs roughly $0.9 per Index task against K3&#39;s $2.3 and runs at about 84 tokens per second against K3&#39;s 41 — K3 is rack-scale to self-host and slow even on Moonshot&#39;s hosted API. Both models have weights genuinely released on Hugging Face (K3 on 27 July, GLM-5.3 on 28 August) and both ship under vendor-specific licences rather than a standard open one, so neither wins on licence. GLM-5.3 is therefore ahead or level on quality and clearly ahead on serving cost.&lt;/p&gt;
&lt;p&gt;Two caveats. GLM-5.3 on Hugging Face is a text-generation model under Z.ai&#39;s custom glm-5.3 licence — check its terms before commercial use, and note it does not match K3&#39;s native vision, though vision is not a criterion here. GLM-5.3-Flash (MIT, 26 August) scores 42 and is the cheaper Flash-class option. DeepSeek-V4.1-Flash appeared on Hugging Face trending on 10 September under MIT with an FP8 checkpoint, but carries no independent score yet.&lt;/p&gt;</description>
  </item>
  <item>
    <title>[CHALLENGED] Wan 3.0 takes nominal #1 on AA text-to-video, but on level points with Gemini Omni Flash</title>
    <link>https://rightsignal.co.uk/slots/video-gen/</link>
    <guid isPermaLink="false">https://rightsignal.co.uk/changelog/2026-09-09-video-gen</guid>
    <pubDate>Wed, 09 Sep 2026 06:00:00 +0000</pubDate>
    <description>&lt;p&gt;Artificial Analysis&#39;s text-to-video leaderboard now lists Alibaba Wan 3.0 first and Gemini Omni Flash second — but both are on 1239 Elo, so this is a tiebreak placing rather than a measured lead. Fal&#39;s post-trained Minimax H3 Max is four points back on 1235, MiniMax H3 Open Weights on 1228 and Dreamina Seedance 2.0 720p on 1222, which keeps the top of the board tightly bunched (&lt;a href=&#34;https://artificialanalysis.ai/video/leaderboard/text-to-video&#34;&gt;Artificial Analysis&lt;/a&gt;).&lt;/p&gt;
&lt;p&gt;We need more than a rank swap on equal points to move a title. Nothing in this week&#39;s evidence touches three of our four criteria: there is still no independent read on Wan&#39;s character consistency or single-pass clip length, and no fresh with-audio result to change the existing picture, where the pick is level with or behind Wan on text-to-video and behind Minimax H3 Max on image-to-video. Native audio remains the softest part of the incumbent&#39;s case, but it is not the part Wan has newly won.&lt;/p&gt;
&lt;p&gt;So Gemini Omni Flash keeps the title, with the same caveats as before: 720p standard output with upscaling as a separate pass, so rivals&#39; &#34;15 seconds at 2K&#34; figures are not like-for-like, and Google&#39;s own model card still lists complete consistency across edits as an open limitation.&lt;/p&gt;
&lt;p&gt;What would settle it: a second independent board putting Wan 3.0 clearly ahead, or the same AA gap opening beyond noise and holding across successive reads — plus any credible third-party measurement of clip length and character consistency, where Wan is currently unmeasured.&lt;/p&gt;</description>
  </item>
  <item>
    <title>[CHANGED] Breeze TTS 2 Open Weights takes best open-weight TTS from Fish Audio S2 Pro</title>
    <link>https://rightsignal.co.uk/slots/tts-open-weight/</link>
    <guid isPermaLink="false">https://rightsignal.co.uk/changelog/2026-09-09-tts-open-weight-2</guid>
    <pubDate>Wed, 09 Sep 2026 06:00:00 +0000</pubDate>
    <description>&lt;p&gt;Fish Audio S2 Pro has held this slot since March, but the Artificial Analysis provider-voice board no longer supports it. In the 9 September standings the pick sits at 1128 Elo in 25th place and is out of the top ten, while BreezeBlue&#39;s Breeze TTS 2 Open Weights is 7th at 1215 and the highest-placed open-weight model on the board. That is an 87-point gap on the slot&#39;s principal quality measure, not a rounding error, and the weights have been public on Hugging Face since 25 August 2026.&lt;/p&gt;
&lt;p&gt;Breeze also wins the remaining criteria. BreezeBlue quotes under 40ms time-to-first-audio and 0.32 RTF on a warmed-up H100, against roughly 100ms first audio and 0.195 RTF on an H200 for S2 Pro, and it needs about 7.7 GiB for eager inference, so a 12GB card will do rather than H200-class kit. On licence there is no gain and no loss: Breeze&#39;s inference code is Apache 2.0 but the checkpoints carry a research-and-non-commercial licence, exactly as restrictive as the Fish Audio Research License. Commercial deployment still needs a separate agreement in both cases.&lt;/p&gt;
&lt;p&gt;Two caveats worth weighing before you swap. Breeze covers English and Chinese only, where S2 Pro spans 80-plus languages; if you need broad language coverage the old pick remains the better tool despite the Elo gap. And Step-Audio-EditX, at 1104 Elo, is behind both but keeps the loosest code licensing, so it stays the fallback for self-hosters who must ship commercially.&lt;/p&gt;</description>
  </item>
  <item>
    <title>[CHALLENGED] Breeze TTS 2 challenges Fish Audio S2 Pro for best open-weight TTS</title>
    <link>https://rightsignal.co.uk/slots/tts-open-weight/</link>
    <guid isPermaLink="false">https://rightsignal.co.uk/changelog/2026-09-09-tts-open-weight</guid>
    <pubDate>Wed, 09 Sep 2026 06:00:00 +0000</pubDate>
    <description>&lt;p&gt;Breeze TTS 2 would take this title on today&#39;s evidence, but the case has not yet cleared our bar for a change, so it is recorded as the challenger for now. The adjudicator&#39;s reasoning:&lt;/p&gt;
&lt;p&gt;The Artificial Analysis provider-voice arena has settled the question. BreezeBlue Breeze TTS 2 Open Weights now sits at #7 overall on 1215 Elo — the highest-placed open-weights model on the board — while Fish Audio S2 Pro has fallen out of the top ten altogether. That is not a one-off reading: the same board had Breeze at 1,220 against 1,125 for S2 Pro shortly after the weights went up on Hugging Face on 25 August 2026, and the gap has held through today&#39;s fetch.&lt;/p&gt;
&lt;p&gt;Breeze also takes the two engineering criteria. BreezeBlue quotes under 40ms time-to-first-audio on an H100 with roughly 7.7GiB for eager inference, a 12GB GPU minimum and 24GB for the fast path; S2 Pro&#39;s quoted ~100ms first audio and 0.195 RTF assume an H200-class card. On licence there is nothing to choose between them: Breeze ships Apache-2.0 inference code but puts weights, derivatives and self-hosted outputs under the BreezeBlue Research and Non-Commercial Licence, needing written authorisation from RESONIA — the same blocker as the Fish Audio Research License.&lt;/p&gt;
&lt;p&gt;Two caveats stay on the record. Breeze is English and Chinese only, a sharp narrowing from S2 Pro&#39;s 80+ languages; if you need broad multilingual coverage, the old pick is still the one to reach for. And if you need to ship commercially from self-hosted weights, neither of these helps you — Step-Audio-EditX, at 1104 Elo and Apache-2.0 on the repository code (though with no licence file on the weights repo), remains the least restrictive on paper and runs in 12GB of VRAM.&lt;/p&gt;</description>
  </item>
  <item>
    <title>[CHALLENGED] Qwen3.8 27B&#39;s lead widens on Artificial Analysis, but its 16GB fit is still unverified</title>
    <link>https://rightsignal.co.uk/slots/llm-small/</link>
    <guid isPermaLink="false">https://rightsignal.co.uk/changelog/2026-09-09-llm-small</guid>
    <pubDate>Wed, 09 Sep 2026 06:00:00 +0000</pubDate>
    <description>&lt;p&gt;Artificial Analysis has re-scored its small open-models board again. As of 9 September, Qwen3.8 27B holds the top three slots at 34 (xhigh), 31 (medium) and 29 (low), with its non-reasoning mode at 22. Gemma 4 26B A4B, this slot&#39;s title holder, sits at #10 on 17. Both sets of numbers have come down from the August scoring (41/35/34 against 26), so the relative picture is unchanged: Qwen3.8 27B is still the quality leader in this size class, and the gap in index points has if anything widened slightly.&lt;/p&gt;
&lt;p&gt;That still does not take the title. This slot is judged on quality per GB, tool-calling under quantisation and running in 16GB of RAM, and Qwen3.8 27B remains without an official quantised release with published file sizes, without independent 4-bit quality-retention figures and without tool-calling-under-quantisation numbers. Its fit on a 16GB machine is therefore unproven, which is a stated criterion rather than a nice-to-have. Gemma 4 26B A4B keeps the title on deployability: Apache-2.0 weights, a 14.2GB dynamic UD-Q4_K_XL QAT build, native function calling.&lt;/p&gt;
&lt;p&gt;The rest of this week&#39;s dossier is noise for this slot — two unrelated arXiv method papers on preference-data filtering and phase geometry in transformers.&lt;/p&gt;
&lt;p&gt;The title stays under review. A verified quantised Qwen3.8 27B build with published sizes and any independent quantised tool-calling evidence would very likely take the slot on the next pass.&lt;/p&gt;</description>
  </item>
  <item>
    <title>[CHALLENGED] GLM-5.3 challenges Kimi K3 for best open-weight LLM</title>
    <link>https://rightsignal.co.uk/slots/llm-open-weight/</link>
    <guid isPermaLink="false">https://rightsignal.co.uk/changelog/2026-09-09-llm-open-weight</guid>
    <pubDate>Wed, 09 Sep 2026 06:00:00 +0000</pubDate>
    <description>&lt;p&gt;GLM-5.3 would take this title on today&#39;s evidence, but the case has not yet cleared our bar for a change, so it is recorded as the challenger for now. The adjudicator&#39;s reasoning:&lt;/p&gt;
&lt;p&gt;Artificial Analysis now lists GLM-5.3 (max) at 45 on its Intelligence Index, one point ahead of Kimi K3 (max) at 44 — the reverse of the 50-vs-49 reading that put K3 in this slot in July (&lt;a href=&#34;https://artificialanalysis.ai/models&#34;&gt;AA models&lt;/a&gt;).&lt;/p&gt;
&lt;p&gt;Treat the quality gap as a tie: one point, in opposite directions across the two readings, is inside the noise. What settles it is the rest of the slot&#39;s criteria. GLM-5.3&#39;s full weights went public on Hugging Face between 25 and 28 August, so the availability objection that kept it as a challenger is gone. On cost and speed it is not close: roughly $0.9 per Index task against K3&#39;s $2.3, and about 84 tokens per second against K3&#39;s 41. K3 keeps the bigger context window and native vision, and it remains a serious model, but paying more than twice as much for half the throughput to score level is no longer defensible.&lt;/p&gt;
&lt;p&gt;Licence is a wash rather than a win. GLM-5.3 is tagged &#39;other&#39; on Hugging Face, not MIT; K3 shipped under Moonshot&#39;s own custom terms with attribution and MaaS conditions attached. Both need reading before commercial deployment.&lt;/p&gt;
&lt;p&gt;The other open-weight contenders stay behind. DeepSeek-V4-Flash-Vision-Exp scores 42 and is an explicitly experimental Flash-class checkpoint, though its ~168GB FP8/FP4 footprint makes it much the easiest to self-host. GLM-5.3-Flash also sits at 42, MIT-licensed, and is the sensible pick if you need something small and permissive.&lt;/p&gt;
&lt;p&gt;Self-hosting GLM-5.3 is still rack-scale work: reported local serving needs 8x H200 or 10-12x H100 at FP8.&lt;/p&gt;</description>
  </item>
  <item>
    <title>[CHALLENGED] Claude Opus 5 logged as challenger: leads the independent SWE-bench Verified aggregate, where GPT-6 Astra has no run</title>
    <link>https://rightsignal.co.uk/slots/llm-coding/</link>
    <guid isPermaLink="false">https://rightsignal.co.uk/changelog/2026-09-09-llm-coding-3</guid>
    <pubDate>Wed, 09 Sep 2026 06:00:00 +0000</pubDate>
    <description>&lt;p&gt;GPT-6 Astra keeps the title, but with a named challenger.&lt;/p&gt;
&lt;p&gt;The new evidence is benchlm&#39;s independent SWE-bench Verified aggregate, refreshed 9 September, which has Claude Opus 5 first at 96%, ahead of Claude Mythos 5 (95.5%) and Claude Fable 5 (95%). GPT-6 Astra does not appear in that top ten at all — it has no independent SWE-bench Verified run, which was already flagged in the pick&#39;s rationale.&lt;/p&gt;
&lt;p&gt;That is not enough to move the title. Where the two models are directly compared, Astra is still ahead: it is first on the official Terminal-Bench board (v4.0, 58.2%) against Opus 5&#39;s 51.8% in third, and its max tier is first on Arena WebDev at 1796 against claude-opus-5-max&#39;s 1688. Those two boards carry the slot&#39;s agentic and tool-calling criteria between them. A single board on which the pick is simply absent is a gap in the record, not a defeat.&lt;/p&gt;
&lt;p&gt;It is worth logging as a challenge rather than ignoring, because SWE-bench Verified is the closest thing on our boards to repo-scale performance, and it is the one criterion where we have no reading for the pick at all. Two things would settle it: an independent SWE-bench Verified run for Astra, or a second board putting Opus 5 ahead on agentic or tool-calling work. Note also that vals.ai removed SWE-bench Verified from its coding index in August as saturated, so a 96% v 97% ordering there should not be over-read either way.&lt;/p&gt;
&lt;p&gt;Claude Fable 5.1 sits second on both Terminal-Bench 4.0 (57.9%) and Arena WebDev (1764) and remains close behind.&lt;/p&gt;</description>
  </item>
  <item>
    <title>[CHANGED] GPT-6 Astra takes best LLM for coding from Claude Opus 5</title>
    <link>https://rightsignal.co.uk/slots/llm-coding/</link>
    <guid isPermaLink="false">https://rightsignal.co.uk/changelog/2026-09-09-llm-coding-2</guid>
    <pubDate>Wed, 09 Sep 2026 06:00:00 +0000</pubDate>
    <description>&lt;p&gt;Two independent boards now point the same way. The official Terminal-Bench board — now on version 4.0, a harder task set than the retired 2.1 — has GPT-6 Astra first at 58.2%, with Fable 5.1 second (57.9%) and Opus 5 third at 51.8% — the agentic-coding criterion this slot leans on hardest. Arena WebDev has gpt-6-astra-max first at 1796, ahead of claude-fable-5.1-max (1764) and claude-opus-5-max (1688), which is where Opus 5 now sits at #3. The Terminal-Bench result is not a one-off: vals.ai&#39;s own Terminal-Bench 2.1 table also has GPT-6 Astra on top (87.27%, ahead of GPT-5.6 Sol at 85.77% and Claude Fable 5.1 at 85.02%), and it leads vals&#39; Code Migration table too. The reading has held across successive fetches through early September.&lt;/p&gt;
&lt;p&gt;Where the two Astra variants split the boards, Astra leads the agentic and repo-scale tables while the max tier tops human-preference WebDev, so the title goes to Astra.&lt;/p&gt;
&lt;p&gt;What Opus 5 keeps is SWE-bench Verified: it is still #1 on the independent aggregate at 96%. But vals has dropped that benchmark from its coding index as saturated, and the top three there are within a point of each other. Astra has no independent SWE-bench Verified run yet, so its repo-scale case rests on Code Migration alone — the main reason to hold this at medium confidence rather than high.&lt;/p&gt;
&lt;p&gt;Practical note: Astra costs $10/$50 per million tokens against Opus 5&#39;s $5/$25. For cheap, high-volume repo-scale batch work, Opus 5 remains the better buy.&lt;/p&gt;</description>
  </item>
  <item>
    <title>[CHALLENGED] GPT-6 Astra challenges Claude Opus 5 for best LLM for coding</title>
    <link>https://rightsignal.co.uk/slots/llm-coding/</link>
    <guid isPermaLink="false">https://rightsignal.co.uk/changelog/2026-09-09-llm-coding</guid>
    <pubDate>Wed, 09 Sep 2026 06:00:00 +0000</pubDate>
    <description>&lt;p&gt;GPT-6 Astra would take this title on today&#39;s evidence, but the case has not yet cleared our bar for a change, so it is recorded as the challenger for now. The adjudicator&#39;s reasoning:&lt;/p&gt;
&lt;p&gt;The title moves to GPT-6 Astra Max. On the agentic side, vals.ai&#39;s Terminal-Bench 2.1 now has Astra first at 87.27%, ahead of GPT-5.6 Sol (85.77%), Claude Fable 5.1 (85.02%) and Opus 5 in fourth at 84.64% — and Opus 5&#39;s own figure falls to 81.27% once its server-side fallbacks to Opus 4.8 on refused tasks are counted as failures. Astra also tops vals&#39; Code Migration table, which is the nearest thing on that board to repo-scale work. That is not a one-board reading: Arena WebDev has had Astra at #1 across both fetches in this cycle, 1,796 on 9 September against claude-fable-5.1-max at 1,764 and claude-opus-5-max at 1,688, with our outgoing pick third.&lt;/p&gt;
&lt;p&gt;Opus 5 keeps #1 on the independent SWE-bench Verified aggregate at 96%, but the top three there sit within a point of each other and vals has dropped that benchmark from its coding index as saturated, so it no longer carries the weight it did when Opus 5 took the title in July.&lt;/p&gt;
&lt;p&gt;Claude Fable 5.1 Max is the other contender and leads the Vals Index overall, but it trails Astra on both boards this slot tracks and its refusals on bio- and cyber-adjacent tasks remain an operational cost. Astra is dearer at $10/$50 per million tokens against Opus 5&#39;s $5/$25, and there is still no independent SWE-bench Verified run for it — that is the main reason for medium rather than high confidence. Teams already happy with Opus 5&#39;s cost profile have no urgent reason to switch; new agentic setups should default to Astra.&lt;/p&gt;</description>
  </item>
  <item>
    <title>[CHALLENGED] GPT-Image-2.5-Sunburst tops both Arena boards, putting GPT Image 2&#39;s title under challenge</title>
    <link>https://rightsignal.co.uk/slots/image-gen/</link>
    <guid isPermaLink="false">https://rightsignal.co.uk/changelog/2026-09-09-image-gen</guid>
    <pubDate>Wed, 09 Sep 2026 06:00:00 +0000</pubDate>
    <description>&lt;p&gt;GPT Image 2 keeps the title this week, but its position has weakened materially. OpenAI has shipped ChatGPT Images 2.5, and Arena now ranks gpt-image-2 (medium) third on both boards that matter here: third in text-to-image on 1381, behind GPT-Image-2.5-Sunburst (1421) and Flare (1399), and third in single-image edit on 1461, behind Sunburst (1520) and Flare (1491). Those are the slot&#39;s two principal criteria, and the margins — 40 and 59 Elo — are wider than measurement noise.&lt;/p&gt;
&lt;p&gt;What stops a change today is corroboration. The lead rests on a single board, read once. Artificial Analysis has not yet rated any GPT-Image-2.5 variant: its text-to-image arena still has GPT Image 2 (high) first on 1178, 29 clear of MAI-Image-2.6 (1149) and 51 clear of Reve 2.1 (1127). There is also no verified official model page for GPT-Image-2.5 in front of us, and we will not point a title at an unconfirmed listing.&lt;/p&gt;
&lt;p&gt;The previous challenger, MAI-Image-2.6, is now the lesser story. It still leads Artificial Analysis&#39;s editing arena on 1122 against GPT Image 2 (high) on 1117 — a five-point gap, effectively level — and remains far cheaper at $38.90 per 1,000 images against $211. That is a cost argument, not a quality one.&lt;/p&gt;
&lt;p&gt;If Artificial Analysis rates the 2.5 models, or Arena holds these standings on a second reading, this title moves to Sunburst.&lt;/p&gt;</description>
  </item>
  <item>
    <title>[HELD] Weekly review: all 4 titles re-verified</title>
    <link>https://rightsignal.co.uk/changes/</link>
    <guid isPermaLink="false">https://rightsignal.co.uk/changelog/2026-09-09-held-2</guid>
    <pubDate>Wed, 09 Sep 2026 06:00:00 +0000</pubDate>
    <description>&lt;p&gt;All 4 picks were re-verified against the live leaderboards, repositories and release pages. 10 details moved with the evidence and the affected pick pages have been updated.&lt;/p&gt;
&lt;p&gt;Under heightened review after this pass: Claude Opus 5 (best llm for coding), Kimi K3 (best open-weight llm), Claude Fable 5 (best overall llm). A title only changes through full adjudication.&lt;/p&gt;</description>
  </item>
  <item>
    <title>[HELD] Weekly review: all 6 titles re-verified</title>
    <link>https://rightsignal.co.uk/changes/</link>
    <guid isPermaLink="false">https://rightsignal.co.uk/changelog/2026-09-09-held</guid>
    <pubDate>Wed, 09 Sep 2026 06:00:00 +0000</pubDate>
    <description>&lt;p&gt;All 6 picks were re-verified against the live leaderboards, repositories and release pages. 11 details moved with the evidence and the affected pick pages have been updated.&lt;/p&gt;
&lt;p&gt;Under heightened review after this pass: GPT Image 2 (best image generation), Fish Audio S2 Pro (best open-weight tts). A title only changes through full adjudication.&lt;/p&gt;</description>
  </item>
  <item>
    <title>[CHALLENGED] MiniCPM5-2B logged as a challenger; AA small-open board re-scored again</title>
    <link>https://rightsignal.co.uk/slots/llm-small/</link>
    <guid isPermaLink="false">https://rightsignal.co.uk/changelog/2026-09-08-llm-small</guid>
    <pubDate>Tue, 08 Sep 2026 06:00:00 +0000</pubDate>
    <description>&lt;p&gt;No change to the title. Gemma 4 26B A4B stays because nothing here disturbs the argument that won it: a 14.2GB dynamic 4-bit build that actually loads in 16GB, with native function calling.&lt;/p&gt;
&lt;p&gt;Two things worth recording. First, Artificial Analysis has re-scored the small open board again: Qwen3.8 27B now reads 34 at xhigh effort, 31 medium, 29 low and 22 non-reasoning, down from 52/44/43 in the 22 August snapshot (&lt;a href=&#34;https://artificialanalysis.ai/models/open-source/small&#34;&gt;AA small open&lt;/a&gt;). It remains the quality leader in the class, and the underlying blocker is unchanged — no official quantised release with published file sizes, so its fit in 16GB is still unproven. A new entrant, Qwen3.6 35B A3B (Reasoning), arrives at #5 with 22. The MoE shape is interesting for on-device decode speed, but 35B total gives it no obvious path into 16GB at 4-bit and there is no published sizing to check.&lt;/p&gt;
&lt;p&gt;Second, MiniCPM5-2B is trending on Hugging Face, with the card tagging long context, tool-calling and edge deployment (&lt;a href=&#34;https://huggingface.co/openbmb/MiniCPM5-2B&#34;&gt;model card&lt;/a&gt;). At 2B it would be a quality-per-GB argument rather than a quality one, and right now there is no independent benchmark placement and no quantised file sizes to verify. Logged as a challenger, not a contender.&lt;/p&gt;</description>
  </item>
  <item>
    <title>[CHALLENGED] GPT-6 Astra draws level with Claude Fable 5.1 on the Artificial Analysis Intelligence Index</title>
    <link>https://rightsignal.co.uk/slots/llm-overall/</link>
    <guid isPermaLink="false">https://rightsignal.co.uk/changelog/2026-09-08-llm-overall</guid>
    <pubDate>Tue, 08 Sep 2026 06:00:00 +0000</pubDate>
    <description>&lt;p&gt;Artificial Analysis has re-scored its Intelligence Index leaderboard and the top of the board is now effectively level. The listed #1 is Claude Fable 5.1 (Adaptive Reasoning, Xhigh Effort, Default Fallback) on 53, with GPT-6 Astra (max) at #2 and GPT-6 Astra (xhigh) at #3, both also on 53. Claude Fable 5.1 (High Effort) sits fourth on 51, with GPT-6 Astra (high) fifth on the same score.&lt;/p&gt;
&lt;p&gt;Two things follow. First, the incumbent keeps the slot: on the one independent composite that moved, the Fable line is still ranked first, and nothing in this evidence touches the Arena text board or the official Terminal-Bench 2.1 board, where Fable 5 holds 83.8% with Claude Code at xhigh effort.&lt;/p&gt;
&lt;p&gt;Second, the GPT-6 Astra challenger note is now out of date. It described Astra as sitting two to three points behind on the Index; on the current snapshot it is tied on score and separated only by ordering. That is a genuine narrowing, but a tie on a single composite is not a win on the slot&#39;s criteria, and Astra still has no Arena or Terminal-Bench 2.1 placement to corroborate it.&lt;/p&gt;
&lt;p&gt;Note also that Claude Opus 5, previously the strongest case for a change on this board, is not visible in the current top five. Index scores have clearly been rebased — the absolute numbers here are lower than the 62/63 figures carried in the caveats — so read exact standings off the live board rather than comparing across snapshots.&lt;/p&gt;</description>
  </item>
  <item>
    <title>[CHALLENGED] Microsoft&#39;s VibeVoice-ASR-Streaming-7B logged as an unverified challenger; ARK-ASR-3B holds</title>
    <link>https://rightsignal.co.uk/slots/stt-open-weight/</link>
    <guid isPermaLink="false">https://rightsignal.co.uk/changelog/2026-09-07-stt-open-weight</guid>
    <pubDate>Mon, 07 Sep 2026 06:00:00 +0000</pubDate>
    <description>&lt;p&gt;No change this week. The bulk of the incoming evidence is literature rather than evaluation: papers on ASR hallucination under environmental degradation, a word-level backdoor attack (GhostWord), turn-aware streaming supervision, a trilingual spoken-dialogue fact-checking benchmark, and a Mandarin-English code-switching dataset. Interesting reading, but none of it scores a model against the criteria this slot is judged on, so none of it moves ARK-ASR-3B.&lt;/p&gt;
&lt;p&gt;The one item worth recording is microsoft/VibeVoice-ASR-Streaming-7B, which appeared on Hugging Face&#39;s trending list on 2 September 2026. From the listing we can see it is an automatic-speech-recognition pipeline tagged for streaming and for four languages (en, zh, es, pt). That is the whole of the public signal. There is no independent leaderboard result, no reported throughput, and the listing surfaces no licence — and without a licence you cannot even confirm it belongs in an open-weight slot, let alone whether it beats a 4.76 mean WER at RTFx 490.98.&lt;/p&gt;
&lt;p&gt;A streaming 7B from Microsoft is plausibly relevant to voice-agent workloads where latency, not batch WER, decides the design. But this slot ranks on independent WER first, and until the model is scored with the Open ASR Leaderboard&#39;s own harness there is nothing to compare. It joins the existing challenger list on the same terms as Orze-ASR-3Way: a repo exists, a claim does not. Check the card yourself for licence terms before building against it.&lt;/p&gt;</description>
  </item>
  <item>
    <title>[HELD] Weekly sweep: 10 slots reviewed, no changes</title>
    <link>https://rightsignal.co.uk/changes/</link>
    <guid isPermaLink="false">https://rightsignal.co.uk/changelog/2026-09-07-held</guid>
    <pubDate>Mon, 07 Sep 2026 06:00:00 +0000</pubDate>
    <description>&lt;p&gt;75 candidate event(s) examined across all slots; none met promotion criteria. All last-reviewed dates refreshed.&lt;/p&gt;</description>
  </item>
  <item>
    <title>[CHALLENGED] Qwen3.5 Omni enters the Artificial Analysis STT top five; ARK-ASR-3B holds</title>
    <link>https://rightsignal.co.uk/slots/stt-open-weight/</link>
    <guid isPermaLink="false">https://rightsignal.co.uk/changelog/2026-09-06-stt-open-weight</guid>
    <pubDate>Sun, 06 Sep 2026 06:00:00 +0000</pubDate>
    <description>&lt;p&gt;Artificial Analysis&#39; speech-to-text board reshuffled on 6 September 2026, with Qwen3.5 Omni Flash listed at #1 and Qwen3.5 Omni Plus at #2, ahead of Nova 2 Pro, Amazon Transcribe and Universal-3 Pro. Three of those five are hosted commercial services, so they are out of scope for this slot; the two Qwen3.5 Omni entries are the only plausible open-weight challengers in the listing.&lt;/p&gt;
&lt;p&gt;They do not clear the bar. The listing gives us a rank and a percentage and nothing else: no licence, no confirmation that weights are downloadable, no throughput figure and no language coverage. The percentages as recorded are also hard to read as error rates — the #1 entry is shown at 13.5% while the #5 entry is shown at 3.1% — so whatever the board is ordering on, it is not a straight word error rate, and we will not restate those numbers as WER.&lt;/p&gt;
&lt;p&gt;Crucially, this slot is judged on the Open ASR Leaderboard harness, where ARK-ASR-3B still holds at 4.76 mean WER with RTFx 490.98 under Apache-2.0. Nothing in today&#39;s evidence scores a Qwen3.5 Omni variant on that harness, so there is no like-for-like comparison to act on.&lt;/p&gt;
&lt;p&gt;ARK-ASR-3B therefore keeps the title. Qwen3.5 Omni goes on the challenger list pending three things: confirmation that the weights are published under a usable licence, an Open ASR Leaderboard entry run with the leaderboard&#39;s own scorer, and a throughput number. If those land and beat 4.76, this becomes a change.&lt;/p&gt;</description>
  </item>
  <item>
    <title>[CHALLENGED] NVIDIA publishes an NVFP4 build of Qwen3.8-Flash-Next; Gemma 4 26B A4B holds the slot</title>
    <link>https://rightsignal.co.uk/slots/llm-small/</link>
    <guid isPermaLink="false">https://rightsignal.co.uk/changelog/2026-09-06-llm-small</guid>
    <pubDate>Sun, 06 Sep 2026 06:00:00 +0000</pubDate>
    <description>&lt;p&gt;The only new item this cycle is an NVIDIA-published NVFP4 quantisation of Qwen3.8-Flash-Next, trending on Hugging Face and tagged image-text-to-text, quantized, FP4 against the Qwen/Qwen3.8-Flash-Next base model. That is worth noting because the standing objection to Flash-Next in this slot has been the absence of an official quantised release, but it does not resolve it. NVFP4 via NVIDIA&#39;s Model Optimizer is a datacentre GPU format, not a laptop-friendly GGUF, and the listing carries no published file size, so the 16GB fit remains untested. There is still no independent evaluation of Flash-Next at any precision, and nothing at all on tool-calling retention once quantised — the two criteria that decide this slot.&lt;/p&gt;
&lt;p&gt;The other item in the dossier is the 22 August Artificial Analysis open small-model board, which is already reflected in the current caveats: Qwen3.8 27B leads on intelligence, with Gemma 4 26B A4B the board pick on deployability grounds. Nothing there is new.&lt;/p&gt;
&lt;p&gt;So the title stays with Gemma 4 26B A4B. The argument for it is unchanged and narrow: a 14.2GB QAT int4 GGUF that actually loads on a 16GB machine, Apache-2.0 weights, and native function calling. The moment someone publishes an independently measured 4-bit build of a Qwen3.8-class model that fits the same envelope and holds up on tool calls, this slot should turn over. A vendor FP4 upload is not that.&lt;/p&gt;</description>
  </item>
  <item>
    <title>[CHALLENGED] GPT-6 Astra gets its first independent placement — second, behind the Claude Fable line</title>
    <link>https://rightsignal.co.uk/slots/llm-overall/</link>
    <guid isPermaLink="false">https://rightsignal.co.uk/changelog/2026-09-06-llm-overall</guid>
    <pubDate>Sun, 06 Sep 2026 06:00:00 +0000</pubDate>
    <description>&lt;p&gt;GPT-6 Astra is no longer a vendor-claims-only entry. Artificial Analysis today lists it at #2 on the Intelligence Index (max effort, 55) and #3 (xhigh, 54). That answers the main objection in its previous challenger note, but it places OpenAI&#39;s frontier model behind the incumbent rather than ahead of it: the same board now has Claude Fable 5.1 (adaptive reasoning, max effort) at #1 with 57.&lt;/p&gt;
&lt;p&gt;The other movement is within the Anthropic stack. Claude Opus 5, which had been the strongest case for a change after taking the Intelligence Index lead, has dropped to #4 (54) and #5 (53) on Artificial Analysis, and to #3 on Vellum&#39;s board at 64.7%. Vellum now shows Claude Fable 5.1 first at 65%, with Claude Mythos 5.1 alongside it at 65%. On both independent aggregators, the Claude Fable line has reclaimed the top spot it lost in the summer.&lt;/p&gt;
&lt;p&gt;Two caveats. First, the boards are now listing a 5.1 point release while this slot&#39;s pick is recorded as Claude Fable 5; a same-line version bump is not a title change, but readers should expect the card&#39;s benchmark figures — the Arena text lead and the 83.8% on the official Terminal-Bench 2.1 board — to refer to the earlier build until fresh placements appear. Second, the margins are small: 57 against 55, and 65% against 64.7%. Treat this as the incumbent holding on rather than pulling away, and read the exact standings off the live boards.&lt;/p&gt;</description>
  </item>
  <item>
    <title>[CHALLENGED] gpt-6-astra-max takes #1 on Arena WebDev, but no agentic evidence yet</title>
    <link>https://rightsignal.co.uk/slots/llm-coding/</link>
    <guid isPermaLink="false">https://rightsignal.co.uk/changelog/2026-09-06-llm-coding</guid>
    <pubDate>Sun, 06 Sep 2026 06:00:00 +0000</pubDate>
    <description>&lt;p&gt;A new model, gpt-6-astra-max, has entered the Arena WebDev board at #1 with 1797, ahead of claude-fable-5.1-max at 1762, claude-opus-5-max at 1688, qwen3.8-max-0902 at 1686 and kimi-k3-max at 1674. That is a 35-point margin over the previous leader and a clear gap over the Opus line — larger than the four-point shuffles we have logged on this board over the past few weeks, so it is worth registering rather than filing as noise.&lt;/p&gt;
&lt;p&gt;It does not move the title. Arena WebDev is a human-preference board scoring web front-end output; none of this slot&#39;s criteria — agentic coding benchmarks, repo-scale task performance, tool-calling reliability — are measured by it. We have seen the same pattern twice already: claude-fable-5.1-max and qwen3.8-max-0902 both topped this board without a single independent agentic run appearing afterwards, and neither displaced Opus 5.&lt;/p&gt;
&lt;p&gt;What would change our mind is an independent SWE-bench Verified or Terminal-Bench 2.1 run on a third-party harness. Opus 5 holds the pick on vals.ai&#39;s SWE-bench Verified board at 97.0% and sits second on their Terminal-Bench 2.1 table at 84.64%. Those numbers are already under pressure — GPT-5.6 Sol leads Terminal-Bench 2.1 at 85.77%, DeepSeek V4 Pro is within 0.6 points on SWE-bench Verified, and Opus 5&#39;s terminal figure drops to 81.27% if its server-side fallbacks to Opus 4.8 count as failures. A genuinely strong agentic showing from the new model would likely settle it. Until one is published, Opus 5 stays.&lt;/p&gt;</description>
  </item>
  <item>
    <title>[HELD] Weekly review: all 8 titles re-verified</title>
    <link>https://rightsignal.co.uk/changes/</link>
    <guid isPermaLink="false">https://rightsignal.co.uk/changelog/2026-09-06-held</guid>
    <pubDate>Sun, 06 Sep 2026 06:00:00 +0000</pubDate>
    <description>&lt;p&gt;All 8 picks were re-verified against the live leaderboards, repositories and release pages. 17 details moved with the evidence and the affected pick pages have been updated.&lt;/p&gt;
&lt;p&gt;Under heightened review after this pass: Claude Opus 5 (best llm for coding), Fish Audio S2 Pro (best open-weight tts). A title only changes through full adjudication.&lt;/p&gt;</description>
  </item>
  <item>
    <title>[CHALLENGED] Qwen3.8 27B still leads the small-model board, but its Artificial Analysis scores have been marked down</title>
    <link>https://rightsignal.co.uk/slots/llm-small/</link>
    <guid isPermaLink="false">https://rightsignal.co.uk/changelog/2026-09-05-llm-small</guid>
    <pubDate>Sat, 05 Sep 2026 06:00:00 +0000</pubDate>
    <description>&lt;p&gt;Artificial Analysis&#39;s small open-source board has been re-scored, and the numbers on this slot&#39;s challenger card no longer match it. Qwen3.8 27B now sits at 41 at xhigh effort, 35 at medium and 34 at low, against 52/44/43 when we last recorded the board on 22 August. Qwen3.6 27B (reasoning) has likewise dropped from 38 to 29. A new entry, Qwen3.8 27B in non-reasoning mode, comes in at #5 with 26 — level with Gemma 4 26B A4B.&lt;/p&gt;
&lt;p&gt;The ordering is unchanged: Qwen3.8 27B is still the quality leader in this size class, and still by a clear margin at its higher effort settings. What has changed is the size of that margin, which is now roughly 15 points rather than 26, and the fact that the cheapest configuration of the challenger scores no better than the pick.&lt;/p&gt;
&lt;p&gt;None of this moves the title, because none of it touches the criteria. There is still no official quantised Qwen3.8 27B release with published file sizes, no independent measurement of quality retention at 4-bit, and no tool-calling numbers under quantisation. Until someone publishes a build that demonstrably loads and behaves in 16GB, the 27B dense models remain unproven on the one axis this slot cares about.&lt;/p&gt;
&lt;p&gt;Also in this window: SGLang v0.5.19 added support for Qwen3.8 2.4T-A95B. That is a datacentre model and has no bearing here; we note it only because it will show up in searches for the same family name.&lt;/p&gt;
&lt;p&gt;Gemma 4 26B A4B holds, on its 14.2GB QAT int4 build.&lt;/p&gt;</description>
  </item>
  <item>
    <title>[CHALLENGED] GPT-6 Astra formally launched, but no independent numbers yet</title>
    <link>https://rightsignal.co.uk/slots/llm-overall/</link>
    <guid isPermaLink="false">https://rightsignal.co.uk/changelog/2026-09-05-llm-overall</guid>
    <pubDate>Sat, 05 Sep 2026 06:00:00 +0000</pubDate>
    <description>&lt;p&gt;OpenAI has published a launch post for GPT-6 Astra, describing it as its &#34;most intelligent and aligned model yet&#34; with state-of-the-art capabilities across computer use, coding, cybersecurity and science (&lt;a href=&#34;https://openai.com/index/gpt-6-astra&#34;&gt;openai.com&lt;/a&gt;). That settles the earlier uncertainty over whether the model was real and public: it is now an announced release rather than a rumour circulating on social posts.&lt;/p&gt;
&lt;p&gt;What it does not do is move the title. The only evidence we have is the vendor&#39;s own announcement. There is no Arena text placement, no Artificial Analysis Intelligence Index figure, and no appearance on the official Terminal-Bench 2.1 board — the three independent sources this slot leans on. On our stated criteria, reasoning quality, multimodal capability and long-context reliability are all asserted by OpenAI rather than measured by anyone else. Launch-post benchmark claims are exactly the category of evidence that cannot carry a title here, however plausible they look.&lt;/p&gt;
&lt;p&gt;So Claude Fable 5 holds. It remains top of the Arena text board and first on the official Terminal-Bench 2.1 board at 83.8% with Claude Code at xhigh effort. The standing caveats also still apply: the Arena top band is tight, and Claude Opus 5 leads Artificial Analysis&#39;s Intelligence Index at 63 to Fable&#39;s 62 at roughly a quarter less per Index task, which remains the strongest live case against the incumbent.&lt;/p&gt;
&lt;p&gt;We expect third-party placements for Astra within days to weeks. If it lands at the top of Arena text or Terminal-Bench 2.1, or takes the Intelligence Index outright, this slot will change quickly. Until independent numbers exist, it stays a challenger.&lt;/p&gt;
&lt;p&gt;The vLLM release candidate in this batch is routine and has no bearing on the slot.&lt;/p&gt;</description>
  </item>
  <item>
    <title>[CHALLENGED] GPT-6 Astra reported at the top of Epoch AI&#39;s Capability Index; added as a challenger</title>
    <link>https://rightsignal.co.uk/slots/llm-overall/</link>
    <guid isPermaLink="false">https://rightsignal.co.uk/changelog/2026-09-04-llm-overall</guid>
    <pubDate>Fri, 04 Sep 2026 06:00:00 +0000</pubDate>
    <description>&lt;p&gt;Reported to have raised Epoch AI&#39;s Capability Index record from 163 to 169, but the claim reaches us via social posts rather than the Epoch board, and there is no Arena, Terminal-Bench 2.1 or Artificial Analysis placement yet — access may not even be public.&lt;/p&gt;</description>
  </item>
  <item>
    <title>[CHALLENGED] Wan 3.0 still nominally top on AA text-to-video, but the gap has closed to zero</title>
    <link>https://rightsignal.co.uk/slots/video-gen/</link>
    <guid isPermaLink="false">https://rightsignal.co.uk/changelog/2026-09-03-video-gen</guid>
    <pubDate>Thu, 03 Sep 2026 06:00:00 +0000</pubDate>
    <description>&lt;p&gt;Artificial Analysis&#39;s text-to-video leaderboard updated on 3 September with Alibaba&#39;s Wan 3.0 listed first on 1,238 and Gemini Omni Flash second on 1,238 — the same score. That is a narrowing, not a gain: on 20 August the same board had Wan 3.0 on 1,247 to the pick&#39;s 1,239, an eight-point lead. Fal&#39;s post-trained Minimax H3 Max is third on 1,235, MiniMax H3 Open Weights fourth on 1,227 and ByteDance Seed&#39;s Dreamina Seedance 2.0 720p fifth on 1,221, so the top five are inside eighteen points and the ordering at the head of the table is effectively a coin toss.&lt;/p&gt;
&lt;p&gt;On the slot&#39;s criteria this changes nothing. A tied Elo on one sub-board is not evidence that Wan beats the pick on visual quality, and there is still no independent read on Wan&#39;s character consistency, native audio behaviour or maximum clip length — the three places where Gemini Omni Flash&#39;s multi-turn editing and synced audio on every clip have been the differentiator. The pick also continues to lead AA&#39;s no-audio text-to-video board and arena.ai&#39;s Text-to-Video Arena.&lt;/p&gt;
&lt;p&gt;Wan 3.0 stays the lead challenger and remains the most likely candidate to take this title, but it needs either a durable lead on the with-audio board or a second, independent evaluation covering consistency and clip length. Title holds.&lt;/p&gt;</description>
  </item>
  <item>
    <title>[CHALLENGED] Claude Fable 5.1 Max takes Arena WebDev #1 by a wide margin; logged as challenger</title>
    <link>https://rightsignal.co.uk/slots/llm-coding/</link>
    <guid isPermaLink="false">https://rightsignal.co.uk/changelog/2026-09-03-llm-coding</guid>
    <pubDate>Thu, 03 Sep 2026 06:00:00 +0000</pubDate>
    <description>&lt;p&gt;Anthropic&#39;s claude-fable-5.1-max has entered the Arena WebDev board straight at #1 with 1765 Elo, ahead of qwen3.8-max-0902 (1688), claude-opus-5-max (1687), kimi-k3-max (1674) and qwen3.8-max (1669). Unlike the Qwen3.8-Max result we logged previously — a four-point gap that was a tie in practice — this is a 77-point lead over the next entry and a 78-point lead over the incumbent&#39;s Arena variant, which is not noise.&lt;/p&gt;
&lt;p&gt;It is still not enough to move the title. Arena WebDev is a human-preference board for web front-end work; it does not measure any of this slot&#39;s stated criteria — agentic coding benchmarks, repo-scale task performance, or tool-calling reliability. Fable 5.1 has no independent SWE-bench Verified or Terminal-Bench 2.1 run in the evidence to date, and the earlier Fable 5 terminal figures already came with a disclosed harness caveat (Vals reports Opus 4.8 used as a refusal fallback in both Fable 5&#39;s and Opus 5&#39;s runs).&lt;/p&gt;
&lt;p&gt;Claude Opus 5 therefore holds, on the strength of its independent SWE-bench Verified result and its price position. The existing caveats stand: GPT-5.6 Sol leads Opus 5 on vals.ai&#39;s Terminal-Bench 2.1 table, and DeepSeek V4 Pro is within a point on SWE-bench Verified as an open-weight option. We will revisit if an independent agentic or repo-scale run for Fable 5.1 appears.&lt;/p&gt;
&lt;p&gt;Also noted but not decision-relevant: PaperCompiler, a new arXiv method for repository-level paper-to-code generation, which reports no model ranking bearing on this slot.&lt;/p&gt;</description>
  </item>
  <item>
    <title>[CHALLENGED] Community mixed-precision GGUFs appear for Qwen3.8 27B, but the 16GB question is still open</title>
    <link>https://rightsignal.co.uk/slots/llm-small/</link>
    <guid isPermaLink="false">https://rightsignal.co.uk/changelog/2026-09-02-llm-small</guid>
    <pubDate>Wed, 02 Sep 2026 06:00:00 +0000</pubDate>
    <description>&lt;p&gt;The gap that has kept Qwen3.8 27B in the challenger column is starting to close, but it has not closed. ISTA-DASLab has published a mixed-precision GGUF conversion of Qwen3.8 27B (tagged GSQ/RCO, multimodal, image-text-to-text) which is now trending on Hugging Face (&lt;a href=&#34;https://huggingface.co/ISTA-DASLab/Qwen3.8-27B-GSQ-RCO-GGUF&#34;&gt;model card&lt;/a&gt;). That is the first quantised build of the quality leader in this size class to gain visible traction.&lt;/p&gt;
&lt;p&gt;It does not yet move the title. This slot is decided on quality per GB, tool-calling under quantisation, and whether the thing loads in 16GB of RAM — and none of those are established for this build in the evidence to hand. There are no published file sizes, so the 16GB fit remains unproven; there is no independent measurement of quality retention against the bfloat16 parent, which is precisely the failure mode that made Gemma 4&#39;s QAT int4 build the pick (Unsloth measured 70.2% top-1 for a naive Q4_0 conversion against 85.6% for a dynamic build); and there is no report on function calling at reduced precision.&lt;/p&gt;
&lt;p&gt;The leaderboard picture is unchanged. Artificial Analysis&#39;s small open-source board still shows Qwen3.8 27B holding the top three positions at 52/44/43 against this pick&#39;s 26, with Qwen3.6 27B at 38 and Muse Glimmer at 35 (&lt;a href=&#34;https://artificialanalysis.ai/models/open-source/small&#34;&gt;board&lt;/a&gt;).&lt;/p&gt;
&lt;p&gt;So Gemma 4 26B A4B holds on deployability, not on score. If someone publishes sizes and a quality-retention or tool-calling check for one of these Qwen3.8 quants, expect this title to change.&lt;/p&gt;</description>
  </item>
  <item>
    <title>[CHALLENGED] Qwen3.8-Max-0902 takes #1 on Arena WebDev, joining the challenger list behind Claude Opus 5</title>
    <link>https://rightsignal.co.uk/slots/llm-coding/</link>
    <guid isPermaLink="false">https://rightsignal.co.uk/changelog/2026-09-02-llm-coding</guid>
    <pubDate>Wed, 02 Sep 2026 06:00:00 +0000</pubDate>
    <description>&lt;p&gt;Alibaba&#39;s new checkpoint, qwen3.8-max-0902, has moved to #1 on the Arena WebDev coding board at 1,691 Elo, displacing claude-opus-5-max, which now sits second at 1,687. The rest of the top five is kimi-k3-max at 1,674, the earlier qwen3.8-max at 1,669 and claude-opus-5-high at 1,661 (&lt;a href=&#34;https://arena.ai/leaderboard/code/webdev&#34;&gt;Arena WebDev&lt;/a&gt;).&lt;/p&gt;
&lt;p&gt;A 4-point Elo gap is not a result. On a board of this size that margin is comfortably inside the usual confidence interval, and the two models are best read as tied. More importantly, Arena WebDev measures human preference between generated web front-ends. It tells you little about the three things this slot is judged on: agentic coding benchmarks, repo-scale task completion and tool-calling reliability. Kimi K3 has been sitting in the same neighbourhood on this board for weeks without shifting the title, for the same reason.&lt;/p&gt;
&lt;p&gt;So Qwen3.8-Max goes on the challenger list rather than into the pick. What would move it: an independent agentic run — SWE-bench Verified or Terminal-Bench 2.1 on a third-party harness such as vals.ai — showing it at or above Opus 5&#39;s 97.0% and 84.64%. Nothing in this cycle provides that.&lt;/p&gt;
&lt;p&gt;The existing caveats stand unchanged. Opus 5&#39;s terminal lead remains genuinely contested against GPT-5.6 Sol, and its Terminal-Bench figure still depends on how you score server-side fallbacks to Opus 4.8. Claude Opus 5 holds the title.&lt;/p&gt;</description>
  </item>
  <item>
    <title>[CHALLENGED] Wan 3.0 stretches its with-audio lead over Gemini Omni Flash to eight points</title>
    <link>https://rightsignal.co.uk/slots/video-gen/</link>
    <guid isPermaLink="false">https://rightsignal.co.uk/changelog/2026-08-31-video-gen</guid>
    <pubDate>Mon, 31 Aug 2026 06:00:00 +0000</pubDate>
    <description>&lt;p&gt;Artificial Analysis&#39;s text-to-video board now puts Alibaba Wan 3.0 first on 1247 with Gemini Omni Flash second on 1239, with MiniMax H3 Open Weights third on 1228 and ByteDance&#39;s Dreamina Seedance 2.0 720p fourth on 1221. That widens Wan&#39;s margin over the pick from four points to eight on the same board, and it is the second consecutive read in which Wan sits ahead rather than a one-off.&lt;/p&gt;
&lt;p&gt;It is still not enough to move the title. Eight Elo points on a single crowd-vote sub-board, on a table where AA had lately been publishing a ranked range spanning several positions for Wan, is not a decisive result on visual quality, and the slot is judged on four criteria rather than one. Gemini Omni Flash continues to lead AA&#39;s no-audio text-to-video table and its image-to-video board without audio, and there is still no independent measurement of Wan 3.0&#39;s clip length or character consistency to set against the pick&#39;s multi-turn editing through the Interactions API, which is the practical reason it holds the slot.&lt;/p&gt;
&lt;p&gt;The only other item in this cycle is a community-converted &lt;code&gt;MiniMax-H3-experimental&lt;/code&gt; repo trending on Hugging Face. That is a third-party packaging of a model already tracked as a challenger and carries no evaluation of its own, so it does not shift anything.&lt;/p&gt;
&lt;p&gt;If Wan 3.0&#39;s lead holds or grows once AA&#39;s sample count closes on the pick&#39;s, and any independent read appears on length or consistency, this becomes a change rather than a challenge.&lt;/p&gt;</description>
  </item>
  <item>
    <title>[CHALLENGED] DeepSeek-V4-Flash-Vision-Exp lands under MIT; Kimi K3 keeps the open-weight title</title>
    <link>https://rightsignal.co.uk/slots/llm-open-weight/</link>
    <guid isPermaLink="false">https://rightsignal.co.uk/changelog/2026-08-31-llm-open-weight</guid>
    <pubDate>Mon, 31 Aug 2026 06:00:00 +0000</pubDate>
    <description>&lt;p&gt;An experimental member of the DeepSeek V4 family, &lt;code&gt;deepseek-ai/DeepSeek-V4-Flash-Vision-Exp&lt;/code&gt;, appeared on the Hugging Face trending list on 31 August with safetensors weights, fp8 and 8-bit variants, an image-text-to-text pipeline, eval-results tags and — notably — an MIT licence (&lt;a href=&#34;https://huggingface.co/deepseek-ai/DeepSeek-V4-Flash-Vision-Exp&#34;&gt;model card&lt;/a&gt;). On licence and &#34;weights actually released&#34;, that is stronger than the incumbent: Kimi K3 ships under Moonshot&#39;s own custom terms with attribution and MaaS conditions above certain thresholds.&lt;/p&gt;
&lt;p&gt;It does not take the title. This is a Flash-class, explicitly experimental checkpoint, and nothing in the current evidence gives it an independent head-to-head score against K3 or against the GLM-5.3 weights already logged as a challenger. Flash-tier releases from every lab so far trade quality for serving cost, and serving cost is only one of four criteria here. We would want an Artificial Analysis-style Intelligence Index placement, or another third-party evaluation, on the downloadable weights before treating a V4 checkpoint as the frontier open-weight model.&lt;/p&gt;
&lt;p&gt;Also in the dossier and not moving the needle: &lt;code&gt;orcarouter/Qwen3.8-Flash-Next-Uncensored-GGUF&lt;/code&gt;, a community abliterated quantisation of an existing Flash-class Qwen rather than a new model (&lt;a href=&#34;https://huggingface.co/orcarouter/Qwen3.8-Flash-Next-Uncensored-GGUF&#34;&gt;listing&lt;/a&gt;).&lt;/p&gt;
&lt;p&gt;Kimi K3 therefore holds the slot, still on the strength of being the highest-scoring thing you can download, and still with the same caveats: enormous to self-host and slow in practice. The interesting question for the next few weeks is whether a full, non-experimental DeepSeek V4 arrives under the same MIT terms.&lt;/p&gt;</description>
  </item>
  <item>
    <title>[HELD] Weekly sweep: 10 slots reviewed, no changes</title>
    <link>https://rightsignal.co.uk/changes/</link>
    <guid isPermaLink="false">https://rightsignal.co.uk/changelog/2026-08-31-held</guid>
    <pubDate>Mon, 31 Aug 2026 06:00:00 +0000</pubDate>
    <description>&lt;p&gt;49 candidate event(s) examined across all slots; none met promotion criteria. All last-reviewed dates refreshed.&lt;/p&gt;</description>
  </item>
  <item>
    <title>[HELD] Weekly review: all 10 titles re-verified</title>
    <link>https://rightsignal.co.uk/changes/</link>
    <guid isPermaLink="false">https://rightsignal.co.uk/changelog/2026-08-30-held</guid>
    <pubDate>Sun, 30 Aug 2026 06:00:00 +0000</pubDate>
    <description>&lt;p&gt;All 10 picks were re-verified against the live leaderboards, repositories and release pages. 25 details moved with the evidence and the affected pick pages have been updated.&lt;/p&gt;
&lt;p&gt;Under heightened review after this pass: Kimi K3 (best open-weight llm), Fish Audio S2 Pro (best open-weight tts). A title only changes through full adjudication.&lt;/p&gt;</description>
  </item>
  <item>
    <title>[CHALLENGED] Breeze TTS 2 Open Weights enters the Artificial Analysis top five, challenging Fish Audio S2 Pro</title>
    <link>https://rightsignal.co.uk/slots/tts-open-weight/</link>
    <guid isPermaLink="false">https://rightsignal.co.uk/changelog/2026-08-28-tts-open-weight</guid>
    <pubDate>Fri, 28 Aug 2026 06:00:00 +0000</pubDate>
    <description>&lt;p&gt;The Artificial Analysis TTS leaderboard now lists BreezeBlue Breeze TTS 2 Open Weights at #5 with an Elo of 1,217, behind Cartesia Sonic 3.6 (1,285), SpeechifyAI Simba 3.2 (1,241), Alibaba Qwen-Audio-3.0-TTS-Plus (1,240) and VUI Labs Luna TTS (1,225). Fish Audio S2 Pro, which holds this title partly on the strength of being the highest-ranked open-weights entry on that same board, does not appear in the published top five. On the face of it, an open-weights rival has overtaken the current pick on the arena metric we track.&lt;/p&gt;
&lt;p&gt;That is enough to log a challenger, not to change the title. The leaderboard row tells us nothing about Breeze TTS 2&#39;s actual licence terms — &#34;Open Weights&#34; in a model name is not a verified licence, and this slot cares specifically about whether commercial self-hosting is permitted. We also have no time-to-first-audio figures, no parameter count or VRAM requirement, and no independent confirmation that checkpoints are downloadable and reproduce the arena score. Fish S2 Pro&#39;s current Elo is not stated either, so the size of the gap is inferred from its absence from the top five rather than measured.&lt;/p&gt;
&lt;p&gt;What would settle it: a published licence file, a reproducible local inference path with latency and GPU figures, and a leaderboard snapshot showing both models&#39; scores side by side. Until then Fish Audio S2 Pro keeps the slot, with the standing caveat that its licence remains non-commercial and Step-Audio-EditX stays the Apache-2.0 fallback. The two arXiv items in this cycle — a Sanskrit chant TTS pipeline and the SPAR-K early-exit decoding scheme — do not bear on the title.&lt;/p&gt;</description>
  </item>
  <item>
    <title>[CHALLENGED] GLM-5.3 weights land on Hugging Face, drawing level with Kimi K3 but not past it</title>
    <link>https://rightsignal.co.uk/slots/llm-open-weight/</link>
    <guid isPermaLink="false">https://rightsignal.co.uk/changelog/2026-08-28-llm-open-weight</guid>
    <pubDate>Fri, 28 Aug 2026 06:00:00 +0000</pubDate>
    <description>&lt;p&gt;Z.ai&#39;s full GLM-5.3 is now downloadable: the &lt;code&gt;zai-org/GLM-5.3&lt;/code&gt; repository appeared on Hugging Face&#39;s trending list on 25 August, and its public release was confirmed on 28 August alongside practical serving notes — roughly 10–12x H100 (or 8x H200) for FP8, about 390–430GB at 4-bit/NVFP4, and 230–250GB under aggressive 2-bit quantisation with quality and context trade-offs.&lt;/p&gt;
&lt;p&gt;That closes the gap this slot flagged a month ago, when GLM-5.3 matched Kimi K3&#39;s score of 60 on the Artificial Analysis Intelligence Index but had no released weights. The &#34;weights actually released&#34; criterion is now satisfied. What has not changed is the quality picture: matching K3 is not beating it, and nothing in this cycle&#39;s evidence is an independent head-to-head evaluation putting GLM-5.3 ahead on the slot&#39;s stated criteria. Our bar for a title change is an independent result showing the challenger wins, so K3 keeps the slot.&lt;/p&gt;
&lt;p&gt;Two cautions for anyone planning around this. The Hugging Face model card tags GLM-5.3 as &lt;code&gt;license:other&lt;/code&gt;, not MIT — the MIT terms noted previously applied to the smaller GLM-5.3-Flash checkpoint, and the flagship&#39;s licence should be read directly before commercial use. And the local-hardware figures above come from a single social-media summary, not a measured serving benchmark, so treat them as an order-of-magnitude guide rather than a costed comparison against K3&#39;s 2.8T-parameter, 104B-active footprint.&lt;/p&gt;
&lt;p&gt;We will revisit as soon as a third-party board publishes a separated score, or a like-for-like throughput and cost comparison, for the released GLM-5.3 checkpoint against K3.&lt;/p&gt;</description>
  </item>
  <item>
    <title>[CHALLENGED] Fal&#39;s post-trained MiniMax H3 Max enters AA&#39;s top three, two points off the pick</title>
    <link>https://rightsignal.co.uk/slots/video-gen/</link>
    <guid isPermaLink="false">https://rightsignal.co.uk/changelog/2026-08-27-video-gen</guid>
    <pubDate>Thu, 27 Aug 2026 06:00:00 +0000</pubDate>
    <description>&lt;p&gt;Artificial Analysis&#39;s with-audio text-to-video board has tightened again. As of 27 August the top five reads Alibaba Wan 3.0 on 1,240, Gemini Omni Flash on 1,237, Fal&#39;s post-trained MiniMax H3 Max on 1,235, MiniMax H3 Open Weights on 1,227 and Dreamina Seedance 2.0 720p on 1,221. A week earlier the same board had Wan 3.0 on 1,247 and the pick on 1,239, so the nominal gap at the top has narrowed from eight points to three.&lt;/p&gt;
&lt;p&gt;The new name is Fal MiniMax H3 Max, a post-train of MiniMax&#39;s model by fal, arriving straight into third. Three models now sit inside five Elo points of each other, which on this board is a tie rather than a ranking. There is no independent evidence in this cycle on the other slot criteria — clip length, character consistency, image-to-video — for any of the three, and the pick retains its lead on the no-audio board and on arena.ai.&lt;/p&gt;
&lt;p&gt;Google also published Gemini Omni 1.1 Flash, framed around more build-time control. That is a vendor post with no third-party scores attached, so it does not shift the pick&#39;s standing or its recorded version; we will wait for arena and AA placements before treating 1.1 as the tracked build.&lt;/p&gt;
&lt;p&gt;No change to the title. Wan 3.0 and MiniMax H3 remain the standing challengers, now joined by fal&#39;s post-train, and the audio-track weakness flagged in the caveats is unresolved.&lt;/p&gt;</description>
  </item>
  <item>
    <title>[CHALLENGED] Qwen3.8-Flash-Next arrives with day-two GGUFs but no independent scores yet</title>
    <link>https://rightsignal.co.uk/slots/llm-small/</link>
    <guid isPermaLink="false">https://rightsignal.co.uk/changelog/2026-08-27-llm-small</guid>
    <pubDate>Thu, 27 Aug 2026 06:00:00 +0000</pubDate>
    <description>&lt;p&gt;A new Qwen entry, &lt;strong&gt;Qwen3.8-Flash-Next&lt;/strong&gt;, went up on Hugging Face on 24 August and is trending, with an unsloth GGUF conversion following on 26 August (&lt;a href=&#34;https://huggingface.co/Qwen/Qwen3.8-Flash-Next&#34;&gt;model&lt;/a&gt;, &lt;a href=&#34;https://huggingface.co/unsloth/Qwen3.8-Flash-Next-GGUF&#34;&gt;GGUFs&lt;/a&gt;). Both cards list it as an image-text-to-text model, so it matches Gemma 4 26B A4B&#39;s multimodal input, and the existence of community quantisations within two days means the practical questions for this slot — file size at 4-bit, headroom for KV cache, whether function calling survives quantisation — are at least testable now rather than hypothetical.&lt;/p&gt;
&lt;p&gt;That is as far as it goes. There is no independent evaluation of Flash-Next in evidence: it does not appear on Artificial Analysis&#39;s small open-source board as of the 22 August snapshot, which still shows Qwen3.8 27B at 52/44/43 across reasoning efforts, Qwen3.6 27B (Reasoning) at 38, Muse Glimmer at 35 and this pick at 26 (&lt;a href=&#34;https://artificialanalysis.ai/models/open-source/small&#34;&gt;AA&lt;/a&gt;). No file sizes are published in the material we have, so the 16GB fit is unverified, and the Hugging Face card records the licence only as &#34;other&#34; — a step down in clarity from the incumbent&#39;s Apache-2.0.&lt;/p&gt;
&lt;p&gt;The incumbent keeps the title on the same grounds as before: official Q4_0 QAT builds at 14.6GB that actually load on a 16GB machine, with native tool calling. Gemma 4 26B A4B remains the weakest of the field on raw index score, and this slot is under active review — but a trending upload with no third-party numbers does not move it.&lt;/p&gt;</description>
  </item>
  <item>
    <title>[CHALLENGED] GLM-5.3-Flash lands under MIT, but no independent scores yet</title>
    <link>https://rightsignal.co.uk/slots/llm-open-weight/</link>
    <guid isPermaLink="false">https://rightsignal.co.uk/changelog/2026-08-27-llm-open-weight</guid>
    <pubDate>Thu, 27 Aug 2026 06:00:00 +0000</pubDate>
    <description>&lt;p&gt;Two smaller open-weight releases turned up on Hugging Face this week: Z.ai&#39;s &lt;a href=&#34;https://huggingface.co/zai-org/GLM-5.3-Flash&#34;&gt;GLM-5.3-Flash&lt;/a&gt;, tagged MIT, and Alibaba&#39;s &lt;a href=&#34;https://huggingface.co/Qwen/Qwen3.8-Flash-Next&#34;&gt;Qwen3.8-Flash-Next&lt;/a&gt;, tagged &lt;code&gt;license:other&lt;/code&gt;. Both picked up community GGUF conversions from Unsloth within a day or two (&lt;a href=&#34;https://huggingface.co/unsloth/GLM-5.3-Flash-GGUF&#34;&gt;GLM&lt;/a&gt;, &lt;a href=&#34;https://huggingface.co/unsloth/Qwen3.8-Flash-Next-GGUF&#34;&gt;Qwen&lt;/a&gt;), which is the usual signal that people are actually running them locally.&lt;/p&gt;
&lt;p&gt;Neither displaces Kimi K3. These are Flash-class checkpoints, and the dossier contains no independent evaluation placing either above K3 on the criteria this slot is judged against. What is notable is the licence: GLM-5.3-Flash ships as MIT, against K3&#39;s custom Moonshot terms with their MAU and revenue thresholds. If the full GLM-5.3 weights follow on the same licence — the release was previously expected around mid-August — that becomes a serious challenge on licence and serving cost simultaneously, assuming the quality holds. On the evidence here it is a Flash variant only, so it goes on the board as a challenger rather than a contender.&lt;/p&gt;
&lt;p&gt;Meanwhile the incumbent&#39;s position on serving cost has quietly improved. &lt;a href=&#34;https://github.com/vllm-project/vllm/releases/tag/v0.28.0&#34;&gt;vLLM v0.28.0&lt;/a&gt; headlines a Kimi-K3 optimisation push including Decode Context Parallel support and fused FlashKDA decode and prefill kernels. That does not fix K3 being enormous to host, but it narrows the practical gap that made the pick uncomfortable.&lt;/p&gt;
&lt;p&gt;No change. Reviewing again when a third-party score for any GLM-5.3 checkpoint appears.&lt;/p&gt;</description>
  </item>
  <item>
    <title>[CHALLENGED] MAI-Image-2.6 (preview) closes to 20 points of GPT Image 2 on Artificial Analysis</title>
    <link>https://rightsignal.co.uk/slots/image-gen/</link>
    <guid isPermaLink="false">https://rightsignal.co.uk/changelog/2026-08-26-image-gen</guid>
    <pubDate>Wed, 26 Aug 2026 06:00:00 +0000</pubDate>
    <description>&lt;p&gt;Artificial Analysis&#39;s text-to-image arena has re-ordered its top five, and the margin behind GPT Image 2 has tightened. The board now reads: GPT Image 2 (high) first at 1371, Microsoft AI&#39;s MAI-Image-2.6-Preview second at 1351, Reve 2.1 third at 1322, Google&#39;s Nano Banana 2 (Gemini 3.1 Flash Image Preview) fourth at 1321, and GPT Image 1.5 (high) fifth at 1310.&lt;/p&gt;
&lt;p&gt;That puts 20 points between the pick and its nearest rival, down from the roughly 48-point cushion recorded at the last review, when Reve 2.1 was the closest challenger on this board. MAI-Image-2.6 has also moved past Reve into second on Artificial Analysis, matching the position it already held on Arena&#39;s text-to-image board.&lt;/p&gt;
&lt;p&gt;No change to the title. The pick still leads the board, and nothing in this evidence shows the challenger ahead on photoreal quality, editability, text rendering or instruction following — the leaderboard movement is a shrinking deficit, not a lead. MAI-Image-2.6 also remains a preview release, so its scores and availability should be treated as provisional.&lt;/p&gt;
&lt;p&gt;The existing caveat still stands: editing is where GPT Image 2 is weakest relative to the field, and the arrival of a stronger Microsoft entrant on the text-to-image side makes it worth watching whether the same model displaces MAI-Image-2.5-Pro at the top of the editing arena. If the text-to-image gap closes further on a settled, non-preview release, this slot is live.&lt;/p&gt;</description>
  </item>
  <item>
    <title>[HELD] Weekly sweep: 10 slots reviewed, no changes</title>
    <link>https://rightsignal.co.uk/changes/</link>
    <guid isPermaLink="false">https://rightsignal.co.uk/changelog/2026-08-24-held</guid>
    <pubDate>Mon, 24 Aug 2026 06:00:00 +0000</pubDate>
    <description>&lt;p&gt;58 candidate event(s) examined across all slots; none met promotion criteria. All last-reviewed dates refreshed.&lt;/p&gt;</description>
  </item>
  <item>
    <title>[CHALLENGED] Orze-ASR-3Way takes #1 on Open ASR; ARK-ASR-3B holds the title pending licence checks</title>
    <link>https://rightsignal.co.uk/slots/stt-open-weight/</link>
    <guid isPermaLink="false">https://rightsignal.co.uk/changelog/2026-08-23-stt-open-weight</guid>
    <pubDate>Sun, 23 Aug 2026 06:00:00 +0000</pubDate>
    <description>&lt;p&gt;The Open ASR leaderboard has a new leader: bosonai/Orze-ASR-3Way at 3.81 mean WER, ahead of ARK-ASR-3B on 4.76, MOSS-Transcribe-preview-2B on 4.87, MOSS-Transcribe-Diarize on 5.17 and Cohere Transcribe on 5.42 (&lt;a href=&#34;https://huggingface.co/datasets/hf-audio/open-asr-leaderboard&#34;&gt;Open ASR&lt;/a&gt;). That is close to a full WER point over the current pick, measured by an independent scorer rather than a vendor card, so it is a serious challenge on the slot&#39;s first criterion.&lt;/p&gt;
&lt;p&gt;It is not yet enough to take the title. This slot judges licence and language coverage alongside WER, and we have no confirmation of Orze&#39;s licence terms, weight release or language list, nor any throughput figure. ARK&#39;s case here was never a single number — it was Apache-2.0 plus 19 languages plus competitive accuracy — and swapping it out on a leaderboard row alone would be the sort of move this tracker exists to avoid. We will revisit once the model card and an independent throughput run are available.&lt;/p&gt;
&lt;p&gt;One useful correction in the same update: ARK-ASR-3B now appears on the leaderboard in its own right at 4.76, which retires the standing caveat that its numbers were card-derived rather than board-verified. The figure sits between the card&#39;s claimed 5.04% and AutoArk&#39;s rerun at 5.13%, and slightly ahead of both.&lt;/p&gt;
&lt;p&gt;The Artificial Analysis speech-to-text movement this week — Fun-Realtime-ASR-preview at 1.7%, Scribe v2, the Azure MAI-Transcribe pair and Smallest AI Pulse Pro (&lt;a href=&#34;https://artificialanalysis.ai/speech-to-text&#34;&gt;AA&lt;/a&gt;) — is all commercial API territory and does not bear on an open-weight slot.&lt;/p&gt;</description>
  </item>
  <item>
    <title>[CHALLENGED] Claude Opus 5 tops Vellum&#39;s leaderboard, tightening the pressure on Fable 5</title>
    <link>https://rightsignal.co.uk/slots/llm-overall/</link>
    <guid isPermaLink="false">https://rightsignal.co.uk/changelog/2026-08-23-llm-overall</guid>
    <pubDate>Sun, 23 Aug 2026 06:00:00 +0000</pubDate>
    <description>&lt;p&gt;Vellum&#39;s public LLM leaderboard, refreshed on 16 August, now puts &lt;strong&gt;Claude Opus 5&lt;/strong&gt; in first place with a composite score of 64.7, ahead of Claude Mythos 5 (64.5), Claude Opus 4.8 (57.9), Claude Sonnet 5 (57.4) and Kimi K3 (56.0). Claude Fable 5, the current title holder, does not appear in that top five.&lt;/p&gt;
&lt;p&gt;That is a genuine data point for Opus 5, and it lines up with the picture we already recorded: Artificial Analysis scores the two as effectively tied (63 against 62) at half the list price, and Opus 5 sits inside the chasing cluster on Arena preference voting.&lt;/p&gt;
&lt;p&gt;It is not yet enough to move the title. Fable 5&#39;s absence from Vellum&#39;s top five is ambiguous — the board does not tell us whether it was evaluated and fell short or simply is not covered — and a single composite leaderboard cannot outweigh Fable&#39;s continued first place on Arena text and on the official Terminal-Bench 2.1 board (83.8% with Claude Code at xhigh effort). The slot&#39;s criteria weight long-context reliability and multimodal capability too, neither of which this evidence speaks to.&lt;/p&gt;
&lt;p&gt;The practical reading for anyone choosing today is unchanged but sharpening: Fable 5 remains first-among-equals at the top of a statistical tie, while Opus 5 is now the pick that at least one independent board ranks above every other Anthropic model. If a second independent evaluation shows Opus 5 clearly ahead of Fable on reasoning or long-context work, this title changes hands.&lt;/p&gt;</description>
  </item>
  <item>
    <title>[HELD] Weekly review: all 9 titles re-verified</title>
    <link>https://rightsignal.co.uk/changes/</link>
    <guid isPermaLink="false">https://rightsignal.co.uk/changelog/2026-08-23-held</guid>
    <pubDate>Sun, 23 Aug 2026 06:00:00 +0000</pubDate>
    <description>&lt;p&gt;All 9 picks were re-verified against the live leaderboards, repositories and release pages. 19 details moved with the evidence and the affected pick pages have been updated.&lt;/p&gt;</description>
  </item>
</channel>
</rss>