Gemma 4 26B A4B
As of 14 Sep 2026, Gemma 4 26B A4B is the RightSignal pick for best small / on-device LLM.
A 25.2B-total MoE with only 3.8B active parameters, so it decodes at small-model speed, and Google's quantisation-aware training keeps 4-bit quality close to bfloat16.
OpenBMB's official GGUF build is now up, tagged for tool-calling and on-device use, and the file sizes are published — 1.56GB at Q4_K_M, 2.68GB at Q8_0, 5.04GB at F16 — so it clears a 16GB budget with room to spare. What is still missing is independent 4-bit retention data and any tool-calling-under-quantisation numbers. On Artificial Analysis's re-based v4.3 index it sits at 13 against this pick's 17, so it is a deployability story rather than a quality one.
Why it’s the challengercurrent reign
held off
reigns
Why Gemma 4 26B A4B is the pick
- A 25.2B-total MoE with only 3.8B active parameters, so it decodes at small-model speed, and Google's quantisation-aware training keeps 4-bit quality close to bfloat16.
- Be specific about which file you pull: Unsloth measured a naive Q4_0 conversion of the QAT checkpoint at 70.2% top-1 against 85.6% for their dynamic UD-Q4_K_XL build at 14.2GB, so prefer that over Google's own 14.4GB q4_0 GGUF, and budget the roughly 15GB of total memory it wants — it loads on a 16GB machine but leaves little room for KV cache, so plan for modest context rather than the full 256K.
- The pick is about deployability: Apache-2.0 weights, QAT builds sized for laptops, native function calling and image input, rather than raw benchmark score.
Evidence
The sources behind this title’s record.
- huggingface.co/openbmb/MiniCPM5-2B-GGUF
- artificialanalysis.ai/models/open-source/small
- huggingface.co/google/gemma-4-26B-A4B-it
- arxiv.org/abs/2609.05899
- arxiv.org/abs/2609.06000
- huggingface.co/openbmb/MiniCPM5-2B
- huggingface.co/nvidia/Qwen3.8-Flash-Next-NVFP4
- github.com/sgl-project/sglang/releases/tag/v0.5.19
Vendor numbers are treated as claims until independently reproduced — how we judge. Structured benchmark comparisons are on the roadmap.
Caveats & challengers
- Qwen3.8 27B is still the quality leader in this size class — on Artificial Analysis's re-based v4.3 index it scores 34 at xhigh (28 medium, 26 low) against this pick's 17, with the non-reasoning variant at 22 — but the only official quantised build is Qwen's FP8 repo, so a 16GB fit remains unproven, and the third-party 4-bit builds that do exist, such as Unsloth's 23.4GB NVFP4, are GPU-oriented and well over a 16GB budget.
- The pick's QAT int4 GGUF is 14.2GB in Unsloth's dynamic build, about 200MB smaller than a naive Q4_0 conversion of the same checkpoint, and wants roughly 15GB of total memory; that gap is the whole argument for it.
- The title stays under review because a verified 4-bit Qwen3.8 build with published file sizes would likely take the slot.
OpenBMB's official GGUF build is now up, tagged for tool-calling and on-device use, and the file sizes are published — 1.56GB at Q4_K_M, 2.68GB at Q8_0, 5.04GB at F16 — so it clears a 16GB budget with room to spare. What is still missing is independent 4-bit retention data and any tool-calling-under-quantisation numbers. On Artificial Analysis's re-based v4.3 index it sits at 13 against this pick's 17, so it is a deployability story rather than a quality one.
Still the quality leader in this size class — Artificial Analysis's re-based v4.3 index puts it at 34/28/26 across xhigh/medium/low against this pick's 17, with non-reasoning at 22. Qwen's only official quantised build remains the FP8 release, a GPU-oriented repo; there is no official 4-bit or GGUF build and nothing published on tool-calling under quantisation. Third-party NVFP4 checkpoints have now appeared with head-to-head comparisons, but Unsloth's is 23.4GB and aimed at datacentre GPUs, so the 16GB fit stays unverified.
NVIDIA has now published an NVFP4 quantisation on Hugging Face, but it is a GPU-oriented format with no listed file size, no independent benchmark placement and no tool-calling-under-quantisation numbers, so the 16GB fit is still unverified; the base card also lists its licence only as "other".
At a glance
- Title
- Best small / on-device LLM
- Licence
- Apache-2.0
- Pick since
- 2 Apr 2026
- Last reviewed
- 13 Sep 2026
- Title holders to date
- 10
- Official page
- huggingface.co/google/gemma-4-26B-A4B-it
Title history
Every change, on the record.
A 25.2B-total MoE with only 3.8B active parameters, so it decodes at small-model speed, and Google's quantisation-aware training keeps 4-bit quality close to bfloat16.
165 daysas pick
A dense 9B that finally beat gpt-oss-20b outright while leaving far more headroom for context on a 16GB machine.
31 daysas pick
21B total with 3.6B active and native MXFP4, engineered by OpenAI to run in exactly 16GB — the first model where on-device was the design…
209 daysas pick
Added switchable thinking mode at a size that still quantises into 16GB, and swept the small-model benchmarks at launch.
98 daysas pick
Brought 128k context and vision to the 16GB tier, with official QAT int4 checkpoints — the quantised build was the intended artefact, not…
48 daysas pick
Synthetic-data training gave this 14B maths and reasoning scores well above its weight class, with MIT weights.
63 daysas pick
14B at 4-bit still fits comfortably in 16GB, and it clearly outscored Llama 3 8B on coding and multilingual work — local consensus moved …
111 daysas pick
A 15T-token training run at 8B put it near Llama 2 70B quality — the obvious daily driver on consumer hardware through mid-2024.
154 daysas pick
Beat Llama 2 13B at roughly half the size under Apache-2.0, and became the base every local finetune was built on for six months.
204 daysas pick
as pick