ARK-ASR-3B
As of 14 Sep 2026, ARK-ASR-3B is the RightSignal pick for best open-weight STT.
ARK-ASR-3B heads the Open ASR Leaderboard's public listing at 4.76 mean WER with an RTFx of 490.98, ahead of MOSS-Transcribe-preview-2B on 4.87, MOSS-Transcribe-Diarize on 5.17 and Cohere Transcribe on 5.42, with per-dataset results filed in the repo and dated 23 June 2026.
Microsoft has now filled in the gaps: the 7B card ships under MIT, and the streaming technical report (arXiv 2609.02812) quotes a real-time factor at or below 0.104 on an A100, with ten languages supported rather than the four originally tagged. The accuracy numbers are still the authors' own — it does not appear on the Open ASR Leaderboard, which ARK-ASR-3B continues to head at 4.76 mean WER. Worth a look if you need streaming speaker-attributed transcription, but you cannot yet compare it like-for-like on WER.
Why it’s the challengercurrent reign
held off
reigns
Why ARK-ASR-3B is the pick
- ARK-ASR-3B heads the Open ASR Leaderboard's public listing at 4.76 mean WER with an RTFx of 490.98, ahead of MOSS-Transcribe-preview-2B on 4.87, MOSS-Transcribe-Diarize on 5.17 and Cohere Transcribe on 5.42, with per-dataset results filed in the repo and dated 23 June 2026.
- The 5.04% on the card is the seven-set average that omits TEDLIUM; add TEDLIUM's 2.79% and you get the board figure.
- It ships Apache-2.0 with 19 languages, but that is not what separates it from MOSS: only the preview-2B is English-only, while MOSS-Transcribe-Diarize is also Apache-2.0 and covers 50+ languages, so ARK's edge there is accuracy rather than licence or reach.
Evidence
The sources behind this title’s record.
- huggingface.co/microsoft/VibeVoice-ASR-Streaming-7B
- huggingface.co/Audio8/ARK-ASR-3B
- arxiv.org/abs/2609.04404
- arxiv.org/abs/2609.04260
- arxiv.org/abs/2609.04225
- artificialanalysis.ai/speech-to-text
- huggingface.co/datasets/hf-audio/open-asr-leaderboard
- huggingface.co/ibm-granite/granite-speech-4.1-2b
Vendor numbers are treated as claims until independently reproduced — how we judge. Structured benchmark comparisons are on the roadmap.
Caveats & challengers
- NVIDIA's Parakeet-TDT-0.6B-v3 is still the sensible choice where throughput decides the invoice — 25 European languages, automatic language detection and a 6.34% average WER, so you are trading roughly 1.5 WER points for the speed — but note it ships under CC BY 4.0, not Apache-2.0.
- ARK's headline is board-verified rather than card-only: the leaderboard records 4.76 mean WER at RTFx 490.98, including TEDLIUM at 2.79%.
- Do budget for less throughput headroom than that suggests, though: AutoArk's own rerun of the seven public splits on 8x RTX 4090, scored with the leaderboard scorer, lands at 5.13% WER and an overall RTFx of 197.
Microsoft has now filled in the gaps: the 7B card ships under MIT, and the streaming technical report (arXiv 2609.02812) quotes a real-time factor at or below 0.104 on an A100, with ten languages supported rather than the four originally tagged. The accuracy numbers are still the authors' own — it does not appear on the Open ASR Leaderboard, which ARK-ASR-3B continues to head at 4.76 mean WER. Worth a look if you need streaming speaker-attributed transcription, but you cannot yet compare it like-for-like on WER.
Top two on the Artificial Analysis STT board as of 6 Sep 2026, but the listing gives no licence, weight availability, throughput or language coverage, its percentages do not read as a WER ordering (#1 at 13.5%, #5 at 3.1%), and neither variant appears on the Open ASR Leaderboard where ARK-ASR-3B holds at 4.76.
Boson AI has now published a bosonai/Orze-ASR-3Way repo, so there is a card to read, but no entry for the model appears in the Open ASR Leaderboard listing, which ARK-ASR-3B still heads at 4.76 mean WER. The circulating 3.81 therefore is not a board result, and until it is scored with the leaderboard's own harness it stays unverified. Check the card directly for licence, throughput and language coverage before planning around it.
At a glance
- Title
- Best open-weight STT
- Licence
- Apache-2.0
- Pick since
- 22 Jun 2026
- Last reviewed
- 13 Sep 2026
- Title holders to date
- 7
- Official page
- huggingface.co/Audio8/ARK-ASR-3B
Title history
Every change, on the record.
ARK-ASR-3B heads the Open ASR Leaderboard's public listing at 4.76 mean WER with an RTFx of 490.98, ahead of MOSS-Transcribe-preview-2B on 4.87, MOSS-Transcribe-Diarize on 5.17 and Cohere Transcribe on 5.42, with per-dataset results filed in the repo and dated 23 June 2026.
84 daysas pick
IBM finally pushed past the long-standing Canary Qwen number at 5.33% WER under Apache-2.0 — and 300k downloads say people actually switc…
54 daysas pick
Bolting an LLM decoder onto the Canary encoder hit 5.63% WER and gave transcription plus summarisation in one pass; nothing open beat it …
286 daysas pick
Topped the leaderboard at 600M parameters with absurd throughput — the obvious pick for bulk transcription rather than benchmark chasing.
77 daysas pick
Took the top of the Open ASR Leaderboard off Whisper at a third of the size — and NVIDIA held that spot for over a year.
449 daysas pick
The DevDay refresh cut errors meaningfully over v2 on most languages and became the drop-in replacement overnight.
92 daysas pick
as pick