right signal
Pick Best open-weight STT

ARK-ASR-3B

As of 14 Sep 2026, ARK-ASR-3B is the RightSignal pick for best open-weight STT.

ARK-ASR-3B heads the Open ASR Leaderboard's public listing at 4.76 mean WER with an RTFx of 490.98, ahead of MOSS-Transcribe-preview-2B on 4.87, MOSS-Transcribe-Diarize on 5.17 and Cohere Transcribe on 5.42, with per-dataset results filed in the repo and dated 23 June 2026.

Current pick
since 22 Jun 2026
84
days as pick
Next best (challenger)
microsoft/VibeVoice-ASR-Streaming-7B

Microsoft has now filled in the gaps: the 7B card ships under MIT, and the streaming technical report (arXiv 2609.02812) quotes a real-time factor at or below 0.104 on an A100, with ten languages supported rather than the four originally tagged. The accuracy numbers are still the authors' own — it does not appear on the Open ASR Leaderboard, which ARK-ASR-3B continues to head at 4.76 mean WER. Worth a look if you need streaming speaker-attributed transcription, but you cannot yet compare it like-for-like on WER.

Why it’s the challenger
Reign history
84days
current reign
3challenges
held off
22 JUN 2026became pick
0previous
reigns
Last reviewed: 13 Sep 2026 Reviews are continuous. This pick can change when the evidence changes.

Why ARK-ASR-3B is the pick

  • ARK-ASR-3B heads the Open ASR Leaderboard's public listing at 4.76 mean WER with an RTFx of 490.98, ahead of MOSS-Transcribe-preview-2B on 4.87, MOSS-Transcribe-Diarize on 5.17 and Cohere Transcribe on 5.42, with per-dataset results filed in the repo and dated 23 June 2026.
  • The 5.04% on the card is the seven-set average that omits TEDLIUM; add TEDLIUM's 2.79% and you get the board figure.
  • It ships Apache-2.0 with 19 languages, but that is not what separates it from MOSS: only the preview-2B is English-only, while MOSS-Transcribe-Diarize is also Apache-2.0 and covers 50+ languages, so ARK's edge there is accuracy rather than licence or reach.
Judged on WER on independent leaderboards · throughput · licence · language coverage

Evidence

The sources behind this title’s record.

Vendor numbers are treated as claims until independently reproduced — how we judge. Structured benchmark comparisons are on the roadmap.

Caveats & challengers

  • NVIDIA's Parakeet-TDT-0.6B-v3 is still the sensible choice where throughput decides the invoice — 25 European languages, automatic language detection and a 6.34% average WER, so you are trading roughly 1.5 WER points for the speed — but note it ships under CC BY 4.0, not Apache-2.0.
  • ARK's headline is board-verified rather than card-only: the leaderboard records 4.76 mean WER at RTFx 490.98, including TEDLIUM at 2.79%.
  • Do budget for less throughput headroom than that suggests, though: AutoArk's own rerun of the seven public splits on 8x RTX 4090, scored with the leaderboard scorer, lands at 5.13% WER and an overall RTFx of 197.
Challenger microsoft/VibeVoice-ASR-Streaming-7B

Microsoft has now filled in the gaps: the 7B card ships under MIT, and the streaming technical report (arXiv 2609.02812) quotes a real-time factor at or below 0.104 on an A100, with ten languages supported rather than the four originally tagged. The accuracy numbers are still the authors' own — it does not appear on the Open ASR Leaderboard, which ARK-ASR-3B continues to head at 4.76 mean WER. Worth a look if you need streaming speaker-attributed transcription, but you cannot yet compare it like-for-like on WER.

Challenger Qwen3.5 Omni Flash / Plus

Top two on the Artificial Analysis STT board as of 6 Sep 2026, but the listing gives no licence, weight availability, throughput or language coverage, its percentages do not read as a WER ordering (#1 at 13.5%, #5 at 3.1%), and neither variant appears on the Open ASR Leaderboard where ARK-ASR-3B holds at 4.76.

Challenger bosonai/Orze-ASR-3Way

Boson AI has now published a bosonai/Orze-ASR-3Way repo, so there is a card to read, but no entry for the model appears in the Open ASR Leaderboard listing, which ARK-ASR-3B still heads at 4.76 mean WER. The circulating 3.81 therefore is not a board result, and until it is scored with the leaderboard's own harness it stays unverified. Check the card directly for licence, throughput and language coverage before planning around it.

At a glance

Title
Best open-weight STT
Licence
Apache-2.0
Pick since
22 Jun 2026
Last reviewed
13 Sep 2026
Title holders to date
7
Official page
huggingface.co/Audio8/ARK-ASR-3B

Title history

Every change, on the record.

View full changelog
Pick 22 Jun 2026 – presentARK-ASR-3B Current

ARK-ASR-3B heads the Open ASR Leaderboard's public listing at 4.76 mean WER with an RTFx of 490.98, ahead of MOSS-Transcribe-preview-2B on 4.87, MOSS-Transcribe-Diarize on 5.17 and Cohere Transcribe on 5.42, with per-dataset results filed in the repo and dated 23 June 2026.

84 days
as pick
Pick 29 Apr 2026 – 22 Jun 2026Granite Speech 4.1 2B

IBM finally pushed past the long-standing Canary Qwen number at 5.33% WER under Apache-2.0 — and 300k downloads say people actually switc…

54 days
as pick
Pick 17 Jul 2025 – 29 Apr 2026Canary Qwen 2.5B

Bolting an LLM decoder onto the Canary encoder hit 5.63% WER and gave transcription plus summarisation in one pass; nothing open beat it …

286 days
as pick
Pick 1 May 2025 – 17 Jul 2025Parakeet TDT 0.6B v2

Topped the leaderboard at 600M parameters with absurd throughput — the obvious pick for bulk transcription rather than benchmark chasing.

77 days
as pick
Pick 7 Feb 2024 – 1 May 2025Canary 1B

Took the top of the Open ASR Leaderboard off Whisper at a third of the size — and NVIDIA held that spot for over a year.

449 days
as pick
Pick 7 Nov 2023 – 7 Feb 2024Whisper large-v3

The DevDay refresh cut errors meaningfully over v2 on most languages and became the drop-in replacement overnight.

92 days
as pick
Pick 1 Aug 2023 – 7 Nov 2023Whisper large-v2 98 days
as pick