right signal
snapshot Frozen 31 Aug 2026

The state of AI, August 2026

The full board as it stood at the end of August 2026 — 10 titles, 0 changes during the month. See the current board →

Language

4 titles
Best LLM for coding

Claude Opus 5

Tops vals.ai's independent SWE-bench Verified run at 97.0%, the best of 82 systems on that board, and sits second on their Terminal-Bench 2.1 table at 84.64%. At $5/$25 per million tokens it is half the price of Claude Fable 5 ($10/$50) for scores within a point or two across the coding suite. Tracker refreshes through mid-August still name it the sensible default for repo-scale agentic work.

Challenger: Claude Fable 5 — Still tops the official Terminal-Bench 2.1 board at 83.8% with Claude Code, ahead of Codex on GPT-5.5 at 83.1%. That run dates from 7 June and no top-five entry post-dates Opus 5's 24 July launch — Opus 5 has no submitted run on that board at all, so its absence is a submission gap rather than a weaker score. Treat the Anthropic terminal figures with care either way: Vals discloses that both Fable 5's and Opus 5's runs on its harness used Opus 4.8 as a refusal fallback.

Genuinely contested on terminal work: vals.ai puts GPT-5.6 Sol first on Terminal-Bench 2.1 at 85.77% against Opus 5's 84.64%, and Opus 5's figure falls to 81.27% if you count its server-side fallbacks to Opus 4.8 on refused tasks as failures. The SWE-bench Verified lead is narrowing too — DeepSeek V4 Pro (0813) now sits second at 96.40% against 97.0%, as an open-weight model at a fraction of the cost per task.

38
days of reign
since 24 Jul 2026
Best open-weight LLM

Kimi K3

A 2.8T-parameter MoE (104B active, 1M-token context, native vision) whose weights landed on Hugging Face on 27 July under Moonshot's own licence, and Artificial Analysis still scores it 60 on the Intelligence Index — two ahead of Qwen3.8-2.4T-A95B and the best score of anything you can actually download today. Z.ai's GLM-5.3 now matches that 60 on the same board, but its weights are held back by about a fortnight, so K3 keeps the slot for the moment. It is enormous to self-host and slow at roughly 37 tokens per second even on Moonshot's own API, though nothing downloadable beats it on quality yet.

Challenger: DeepSeek-V4-Flash-Vision-Exp — MIT-licensed V4-family weights trending on Hugging Face from 31 August with fp8/8-bit variants and native vision, but an explicitly experimental Flash-class checkpoint with no independent benchmark against K3 yet.

Licence requires attribution above 100M MAU or $20M/month revenue, and a separate agreement for large model-as-a-service businesses. The hosted API shipped 16 July — the reign is dated from the weights landing on 27 July.

35
days of reign
since 27 Jul 2026
Best overall LLM

Claude Fable 5

Number one on the Arena (formerly LMArena) text leaderboard through the August snapshots, with a 1M-token context window. It debuted on 9 June, was pulled on 12 June when the US applied export controls to it and Mythos 5, and returned on 1 July after the Commerce Department lifted them on 30 June. It also tops the official Terminal-Bench 2.1 board at 83.8% with Claude Code at xhigh effort, just ahead of GPT-5.5 on Codex at 83.1%.

Challenger: Claude Opus 5 — Vellum publishes per-benchmark charts rather than a composite, and the 64.7 is its Humanity's Last Exam board, which Opus 5 tops ahead of Mythos 5 on 64.5; Fable 5 is missing from that chart alone, since Vellum scores it elsewhere at 95% on SWE-Bench and first on OSWorld at 85%. The stronger argument is Artificial Analysis, where Opus 5 (max) takes the Intelligence Index at 63 to Fable's 62 for $5/$25 against $10/$50, and 26% less per Index task. That is the best case yet for a change, but it is not enough while Fable holds the Arena text top spot and first place on the official Terminal-Bench 2.1 board.

The Arena text board is tight at the top: Fable leads at roughly 1,525 with Claude Opus 4.8, GPT-5.5 Pro and Gemini 3.1 Pro Preview clustered behind it, so treat the lead as first-among-equals and read the exact tie band off the live board. Artificial Analysis now ranks Claude Opus 5 (max and xhigh) first on its Intelligence Index at 63 to Fable's 62 — a one-point gap rather than a declared tie — at $2.34 per Index task against Fable's $3.14. On terminal work Fable still holds first place on the official Terminal-Bench 2.1 board at 83.8% with Claude Code at xhigh, a board Opus 5 has yet to appear on at all.

61
days of reign
since 1 Jul 2026
Best small / on-device LLM

Gemma 4 26B A4B

A 25.2B-total MoE with only 3.8B active parameters, so it decodes at small-model speed, and Google's quantisation-aware training keeps 4-bit quality close to bfloat16 — take the QAT int4 GGUF at 14.2GB rather than a naive Q4_0 conversion, which Unsloth measured at 70.2% top-1 against 85.6% for their dynamic build. At around 15GB of total memory it loads on a 16GB machine but leaves little room for KV cache, so plan for modest context rather than the full 256K. The pick is about deployability: Apache-2.0 weights, QAT builds sized for laptops, native function calling and image input, rather than raw benchmark score.

Challenger: Qwen3.8-Flash-Next — New Qwen release trending on Hugging Face with community GGUFs already available two days after launch, and multimodal like the pick — but no independent benchmark placement yet, no published quantised file sizes to test the 16GB fit, and the card lists its licence only as "other".

Qwen3.8 27B is now the quality leader in this size class — 52 on Artificial Analysis's Intelligence Index at xhigh effort against this pick's 26 — but there is still no official quantised release with published file sizes, so its 16GB fit is unproven; Qwen3.6 27B (38, Apache-2.0, 262k, 27.8B dense) sits behind it and is likewise too heavy at 4-bit for a 16GB machine. The pick's QAT int4 GGUF is 14.2GB and wants roughly 15GB of total memory, and that gap is the whole argument for it. This title is under active review chiefly because Qwen3.8 27B has pushed the ceiling so far above the pick that a verified quantised build would likely take the slot.

151
days of reign
since 2 Apr 2026

Speech

2 titles
Best open-weight STT

ARK-ASR-3B

ARK-ASR-3B heads the Open ASR Leaderboard's public listing at 4.76 mean WER with an RTFx of 490.98, ahead of MOSS-Transcribe-preview-2B on 4.87, MOSS-Transcribe-Diarize on 5.17 and Cohere Transcribe on 5.42, with per-dataset results filed in the repo and dated 23 June 2026. The 5.04% on the card is the seven-set average that omits TEDLIUM; add TEDLIUM's 2.79% and you get the board figure. It ships Apache-2.0 with 19 languages, but that is not what separates it from MOSS: only the preview-2B is English-only, while MOSS-Transcribe-Diarize is also Apache-2.0 and covers 50+ languages, so ARK's edge there is accuracy rather than licence or reach.

Challenger: bosonai/Orze-ASR-3Way — Boson AI has now published a bosonai/Orze-ASR-3Way repo, so there is a card to read, but no entry for the model appears in the Open ASR Leaderboard listing, which ARK-ASR-3B still heads at 4.76 mean WER. The circulating 3.81 therefore is not a board result, and until it is scored with the leaderboard's own harness it stays unverified. Check the card directly for licence, throughput and language coverage before planning around it.

NVIDIA's Parakeet-TDT-0.6B-v3 is still the sensible choice where throughput decides the invoice — 25 European languages, automatic language detection and a 6.34% average WER, so you are trading roughly 1.5 WER points for the speed — but note it ships under CC BY 4.0, not Apache-2.0. ARK's headline is board-verified rather than card-only: the leaderboard records 4.76 mean WER at RTFx 490.98, including TEDLIUM at 2.79%. Do budget for less throughput headroom than that suggests, though: AutoArk's own rerun of the seven public splits on 8x RTX 4090, scored with the leaderboard scorer, lands at 5.13% WER and an overall RTFx of 197.

70
days of reign
since 22 Jun 2026
Best open-weight TTS

Fish Audio S2 Pro

Open-sourced on 9 March 2026: a ~4.4B dual-autoregressive model trained on 10M+ hours across 80+ languages, currently the highest-ranked open-weights model on the Artificial Analysis TTS leaderboard, with a production SGLang streaming engine (~100ms to first audio on one H200).

Challenger: BreezeBlue Breeze TTS 2 Open Weights — Now the #1 open-weights model on the Artificial Analysis provider-voice arena at 1,220 Elo against 1,125 for the pick, with weights published on Hugging Face on 25 August 2026. The inference code is Apache-2.0, but the weights, derivatives and self-hosted outputs fall under the BreezeBlue Research and Non-Commercial Licence, so commercial deployment needs written authorisation from RESONIA — the same blocker as the pick. It is English and Chinese only, and BreezeBlue quotes under 40ms time-to-first-audio on an H100 with roughly 7.7GiB for eager inference (12GB GPU minimum, 24GB for the fast path).

Open-weight but not open-source: the Fish Audio Research License permits research and non-commercial use free of charge, and any commercial deployment needs a separate licence from Fish Audio, which keeps Apache-licensed rivals such as Step-Audio-EditX relevant for self-hosters. The downloadable weights are also a generation behind the hosted product — S2.1 Pro arrived on 23 June 2026 as an API-only model with roughly 70ms time-to-first-audio and a claimed 61% win rate over S2 Pro, no S2.1 checkpoints have appeared on the fishaudio Hugging Face org, and the free s2.1-pro-free tier has only been extended to 31 August 2026. Budget for an H200-class GPU if you want the quoted 0.195 RTF and ~100ms first-audio figures locally.

175
days of reign
since 9 Mar 2026

Create

2 titles
Best image generation

GPT Image 2

Top of both Arena boards: #1 on text-to-image at 1382, 51 clear of MAI-Image-2.6-preview on 1331, and #1 on single-image edit at 1462, 23 ahead of Grok Imagine Image 2.0 (low) on 1439. It also leads Artificial Analysis's text-to-image arena at 1,370, 19 ahead of MAI-Image-2.6-Preview and 48 ahead of Reve 2.1, though on that site's editing arena it has slipped to fourth at 1,259, behind MAI-Image-2.6-Preview (1,284), MAI-Image-2.5-Pro (1,272) and Reve 2.1 (1,260). It has held the generation crown since its 21 April launch, on the strength of thinking-mode reasoning before generation, improved multilingual text rendering and more reliable instruction-following.

Challenger: MAI-Image-2.6 (preview) — Second on Artificial Analysis's text-to-image arena at 1,351, just 19 points behind GPT Image 2 (high) on 1,370, and second on Arena's text-to-image board at 1331 against 1382. It has now also taken first place on Artificial Analysis's image-editing arena at 1,284, ahead of GPT Image 2 (high) on 1,259, and sits third on Arena's single-image edit board at 1417. Still a private preview on Microsoft Foundry with no published API pricing, so treat it as a strong editing option you cannot yet buy.

Editing is where the lead has gone: on Artificial Analysis's image-editing arena GPT Image 2 (high) is fourth at 1,259, behind MAI-Image-2.6-Preview (1,284), MAI-Image-2.5-Pro (1,272) and Reve 2.1 (1,260), even though it still tops Arena's single-image edit board at 1462 from Grok Imagine Image 2.0 (low) at 1439. It is also the dearest of the leaders at $211 per 1,000 images on Artificial Analysis's pricing, against $108.50 for MAI-Image-2.5-Pro, $67 for Nano Banana 2 and $48 for MAI-Image-2.5. Treat the challenger gaps as provisional — MAI-Image-2.6 is still a preview with no published pricing, and Grok Imagine Image 2.0's Arena scores are flagged preliminary.

132
days of reign
since 21 Apr 2026
Best video generation

Gemini Omni Flash

Still the top line on arena.ai's Text-to-Video Arena, though it now appears as two entries: the newer gemini-omni-1.1-flash build leads on 1,515 with the original Omni Flash second on 1,512 and Black Forest Labs' FLUX 3 Video third on a preliminary 1,495. On Artificial Analysis it leads the no-audio text-to-video board on 1,324, twenty-three points clear of MiniMax H3 on 1,301, but audio is the soft spot: Alibaba's Wan 3.0 is first on the with-audio table on 1,241 to the pick's 1,237. The differentiator remains conversational multi-turn editing through the Interactions API, where each change builds on the last while holding character and scene consistency, with native synced audio on every clip.

Challenger: Alibaba Wan 3.0 — Now first on Artificial Analysis's text-to-video board on 1247 to Gemini Omni Flash's 1239, widening its lead from four points to eight, with MiniMax H3 Open Weights third (1228) and Dreamina Seedance 2.0 720p fourth (1221). Still a single sub-board result: the pick leads AA's no-audio text-to-video and its image-to-video board without audio, and there is no independent read on Wan's clip length or character consistency.

A hold under real pressure: Wan 3.0 leads AA's with-audio text-to-video table on 1,241 to the pick's 1,237, and on AA's image-to-video board the pick only leads without audio (1,365) — with audio it has slipped to fourth on 1,179, behind fal's Minimax H3 Max (1,202), Dreamina Seedance 2.0 720p (1,190) and MiniMax H3 (1,185). Clip length is the other limit: MiniMax H3 runs to 15 seconds at 2K, FLUX 3 Video to 20 seconds in HD or FHD with native audio, and Seedance 2.5 does 30 seconds in a single pass. Seedance 2.5 still carries no score on either AA board, though it has now entered arena.ai's Text-to-Video Arena fifth on 1,476, below the pick.

62
days of reign
since 30 Jun 2026

Build

2 titles
Best text embeddings

Nemotron-3-Embed-8B

NVIDIA's launch materials put Nemotron-3-Embed-8B-BF16 at number one on the RTEB multilingual board as of 16 July 2026, with a headline RTEB score of 78.5% and 75.5% on MMTEB Retrieval. The model card confirms 34 evaluated languages, a 32,768-token maximum sequence length and 4,096-dimension vectors that can be sliced to 2,048 or 1,024 and re-normalised, with weights open under OpenMDW-1.1 plus 1B BF16 and 1B NVFP4 checkpoints for cheaper serving. The RTEB multilingual board now lists 30 tasks over 22 languages and mixes open with closed datasets, and the ranking rests on a vendor snapshot, so re-check the live table before quoting it.

Challenger: Gemini Embedding 2 — Google's Gemini Embedding 2 is natively multimodal, mapping text, images, video, audio and documents into one vector space; it went to public preview on 10 March 2026 and reached general availability on 22 April 2026 through the Gemini API, Vertex AI and the Gemini Enterprise Agent Platform. Output defaults to 3,072 dimensions and is Matryoshka-truncatable from 128 up, with 768, 1,536 and 3,072 the recommended sizes, over an 8,192-token context, at list pricing of about $0.20 per million text tokens. It remains API-only, so it is not an option if you need weights you can host yourself.

The multilingual picture is contested, not split across incomparable boards: Microsoft reports Harrier-OSS-v1-27B at 74.3 averaged over the same 131 multilingual MTEB v2 tasks on which it cites the MMTEB leaderboard SOTA of 72.3 — the score Tencent's KaLM-Embedding-Gemma3-12B-2511 holds at 72.32 — and puts Harrier top on Borda count. That ranking is a vendor snapshot taken in March/April 2026 and can move as submissions land, so check the live table before repeating it. Either way, MMTEB task-mean and our pick's RTEB NDCG@10 answer different questions: one aggregates all task types, the other is retrieval only.

46
days of reign
since 16 Jul 2026
Best inference serving

vLLM

The like-for-like Spheron runs on Llama 3.3 70B FP8 on a single H100 (three-engine, March 2026; two-engine, June 2026) put SGLang between about 1% and 5% ahead of vLLM on unique-prompt traffic, which is inside run-to-run noise. Both used vLLM v0.18.0 on the legacy model runner, before Model Runner V2 became the default for all dense models in v0.25.0 on 11 July 2026, so neither reflects what ships by default today. Pick vLLM for breadth, hardware coverage and release discipline rather than for any current throughput number.

Challenger: SGLang — Near-parity overall and faster on agentic/structured workloads, with major production adoption (xAI, AMD, LinkedIn, Cursor).

Less of a two-horse race than it looks: TensorRT-LLM led the same March 2026 Spheron H100 70B run by 8-13% at 1-50 concurrent requests and about 16% at 100, once its engine was compiled. SGLang's edge is prefix-driven rather than general — roughly 29% higher throughput on 8B-class ShareGPT-style traffic, and 37% lower TTFT p50 at 50 concurrent requests on a 70B workload with an 80% shared prefix — but turning on vLLM's --enable-prefix-caching narrows that TTFT gap to about 15-18% for a plain shared system prompt. Measure your own prefix overlap before committing, and re-run on current releases: those runs used vLLM v0.18.0 on the legacy model runner, before Model Runner V2 became the default for dense models in v0.25.0.

1168
days of reign
since 20 Jun 2023