right signal
Pick Best small / on-device LLM

Gemma 4 26B A4B

As of 14 Sep 2026, Gemma 4 26B A4B is the RightSignal pick for best small / on-device LLM.

A 25.2B-total MoE with only 3.8B active parameters, so it decodes at small-model speed, and Google's quantisation-aware training keeps 4-bit quality close to bfloat16.

Current pick
since 2 Apr 2026
165
days as pick
Next best (challenger)
MiniCPM5-2B

OpenBMB's official GGUF build is now up, tagged for tool-calling and on-device use, and the file sizes are published — 1.56GB at Q4_K_M, 2.68GB at Q8_0, 5.04GB at F16 — so it clears a 16GB budget with room to spare. What is still missing is independent 4-bit retention data and any tool-calling-under-quantisation numbers. On Artificial Analysis's re-based v4.3 index it sits at 13 against this pick's 17, so it is a deployability story rather than a quality one.

Why it’s the challenger
Reign history
165days
current reign
10challenges
held off
2 APR 2026became pick
0previous
reigns
Last reviewed: 13 Sep 2026 Reviews are continuous. This pick can change when the evidence changes.

Why Gemma 4 26B A4B is the pick

  • A 25.2B-total MoE with only 3.8B active parameters, so it decodes at small-model speed, and Google's quantisation-aware training keeps 4-bit quality close to bfloat16.
  • Be specific about which file you pull: Unsloth measured a naive Q4_0 conversion of the QAT checkpoint at 70.2% top-1 against 85.6% for their dynamic UD-Q4_K_XL build at 14.2GB, so prefer that over Google's own 14.4GB q4_0 GGUF, and budget the roughly 15GB of total memory it wants — it loads on a 16GB machine but leaves little room for KV cache, so plan for modest context rather than the full 256K.
  • The pick is about deployability: Apache-2.0 weights, QAT builds sized for laptops, native function calling and image input, rather than raw benchmark score.
Judged on quality per GB · tool-calling under quantisation · runs in 16GB RAM

Evidence

The sources behind this title’s record.

Vendor numbers are treated as claims until independently reproduced — how we judge. Structured benchmark comparisons are on the roadmap.

Caveats & challengers

  • Qwen3.8 27B is still the quality leader in this size class — on Artificial Analysis's re-based v4.3 index it scores 34 at xhigh (28 medium, 26 low) against this pick's 17, with the non-reasoning variant at 22 — but the only official quantised build is Qwen's FP8 repo, so a 16GB fit remains unproven, and the third-party 4-bit builds that do exist, such as Unsloth's 23.4GB NVFP4, are GPU-oriented and well over a 16GB budget.
  • The pick's QAT int4 GGUF is 14.2GB in Unsloth's dynamic build, about 200MB smaller than a naive Q4_0 conversion of the same checkpoint, and wants roughly 15GB of total memory; that gap is the whole argument for it.
  • The title stays under review because a verified 4-bit Qwen3.8 build with published file sizes would likely take the slot.
Challenger MiniCPM5-2B

OpenBMB's official GGUF build is now up, tagged for tool-calling and on-device use, and the file sizes are published — 1.56GB at Q4_K_M, 2.68GB at Q8_0, 5.04GB at F16 — so it clears a 16GB budget with room to spare. What is still missing is independent 4-bit retention data and any tool-calling-under-quantisation numbers. On Artificial Analysis's re-based v4.3 index it sits at 13 against this pick's 17, so it is a deployability story rather than a quality one.

Challenger Qwen3.8 27B

Still the quality leader in this size class — Artificial Analysis's re-based v4.3 index puts it at 34/28/26 across xhigh/medium/low against this pick's 17, with non-reasoning at 22. Qwen's only official quantised build remains the FP8 release, a GPU-oriented repo; there is no official 4-bit or GGUF build and nothing published on tool-calling under quantisation. Third-party NVFP4 checkpoints have now appeared with head-to-head comparisons, but Unsloth's is 23.4GB and aimed at datacentre GPUs, so the 16GB fit stays unverified.

Challenger Qwen3.8-Flash-Next

NVIDIA has now published an NVFP4 quantisation on Hugging Face, but it is a GPU-oriented format with no listed file size, no independent benchmark placement and no tool-calling-under-quantisation numbers, so the 16GB fit is still unverified; the base card also lists its licence only as "other".

At a glance

Title
Best small / on-device LLM
Licence
Apache-2.0
Pick since
2 Apr 2026
Last reviewed
13 Sep 2026
Title holders to date
10
Official page
huggingface.co/google/gemma-4-26B-A4B-it

Title history

Every change, on the record.

View full changelog
Pick 2 Apr 2026 – presentGemma 4 26B A4B Current

A 25.2B-total MoE with only 3.8B active parameters, so it decodes at small-model speed, and Google's quantisation-aware training keeps 4-bit quality close to bfloat16.

165 days
as pick
Pick 2 Mar 2026 – 2 Apr 2026Qwen3.5-9B

A dense 9B that finally beat gpt-oss-20b outright while leaving far more headroom for context on a 16GB machine.

31 days
as pick
Pick 5 Aug 2025 – 2 Mar 2026gpt-oss-20b

21B total with 3.6B active and native MXFP4, engineered by OpenAI to run in exactly 16GB — the first model where on-device was the design…

209 days
as pick
Pick 29 Apr 2025 – 5 Aug 2025Qwen3-14B

Added switchable thinking mode at a size that still quantises into 16GB, and swept the small-model benchmarks at launch.

98 days
as pick
Pick 12 Mar 2025 – 29 Apr 2025Gemma 3 12B

Brought 128k context and vision to the 16GB tier, with official QAT int4 checkpoints — the quantised build was the intended artefact, not…

48 days
as pick
Pick 8 Jan 2025 – 12 Mar 2025Phi-4

Synthetic-data training gave this 14B maths and reasoning scores well above its weight class, with MIT weights.

63 days
as pick
Pick 19 Sep 2024 – 8 Jan 2025Qwen2.5-14B-Instruct

14B at 4-bit still fits comfortably in 16GB, and it clearly outscored Llama 3 8B on coding and multilingual work — local consensus moved …

111 days
as pick
Pick 18 Apr 2024 – 19 Sep 2024Llama 3 8B Instruct

A 15T-token training run at 8B put it near Llama 2 70B quality — the obvious daily driver on consumer hardware through mid-2024.

154 days
as pick
Pick 27 Sep 2023 – 18 Apr 2024Mistral 7B Instruct

Beat Llama 2 13B at roughly half the size under Apache-2.0, and became the base every local finetune was built on for six months.

204 days
as pick
Pick 1 Aug 2023 – 27 Sep 2023Llama 2 13B Chat 57 days
as pick