Can I Run Gemma 4 26B-A4B Locally? 16–18 GB 4-bit MoE vs 31B Dense
Short answer: Gemma 4 26B-A4B is a MoE (4B active, 256K, text+image) whose Unsloth 4-bit band is 16–18 GB total memory — not the 31B dense table. A 16 GB card sits on that floor. A used 3090 (24 GB) is the comfortable Q4 cell. 8-bit is 28–30 GB and misses one 24 GB card. LocalRig has not measured tok/s.
What is 26B-A4B?
MoE sibling in Gemma 4, not a 26B dense clone of 31B. Unsloth’s variant table: 26B-A4B, MoE, 256K, text+image, “best speed/quality tradeoff for computer use.” Official family README (via Unsloth GGUF, 2026-09-10): hybrid attention like the rest of Gemma 4. Hub GGUF unsloth/gemma-4-26B-A4B-it-GGUF showed ~10.7M downloads on 2026-09-10 — volume, not a ranking.
Do not mix with E2B/E4B phone SKUs (4–6 GB 4-bit) or 12B Unified. Those are other rows in the same Unsloth table.
Unsloth memory bands
Copied 2026-09-10. Units = RAM + VRAM or unified.
| Variant | 4-bit | 8-bit | BF16 / FP16 |
|---|---|---|---|
| 26B A4B | 16–18 GB | 28–30 GB | 52 GB |
| 31B dense | 17–20 GB | 34–38 GB | 62 GB |
Sampling: temperature 1.0, top_p 0.95, top_k 64. Context 262,144. Thinking via <|think|>. Same caveats as 31B: KV is extra; Unsloth’s rule is total memory should exceed the file.
16 GB vs used 3090
16 GB matches Unsloth’s 4-bit floor; 24 GB is the used-card default. At 16 GB you are in the 16–18 GB band with little spare for a long window. A 3090 holds 4-bit with KV headroom. BF16 (52 GB) is dual-GPU or unified-memory territory, not a single 3090.
Browse used RTX 3090 24GB on eBay →
Check RTX 4060 Ti 16GB on Amazon → if you accept Unsloth’s 4-bit floor and short context.
Rent to compare 16 vs 24 GB before buying: RunPod, Vast.ai, GPUMart.
If the job is “one used 24 GB card, dense daily driver,” LocalRig’s current default remains Qwen 3.8 27B Q4. A4B is the Gemma MoE alternative on that same card class.
Runtime
Unsloth GGUF + llama.cpp is the documented path. MTP (Jun 9, 2026) is optional and not a 3090 LocalRig bench. Ollama: no dated library+hardware+tok/s packet on this page. Do not copy 31B NVFP4 Spark recipes onto A4B GGUF.
Related: Ollama vs llama.cpp vs vLLM.
What 16 GB vs 24 GB actually changes on A4B
Weights vs KV, and 4-bit vs 8-bit. Unsloth’s 4-bit band is 16–18 GB. A 16 GB card is therefore a weights-at-the-floor machine: Q4 can load, then every extra thousand tokens of context and every thinking trace compete with CUDA and the OS. A 24 GB 3090 is the same 4-bit file with roughly a third of the card left for cache. That is why LocalRig’s used-card default is still the 3090 even when 16 GB “matches the table.”
8-bit at 28–30 GB is not a 3090. Dual 3090 (48 GB) is the first 48 GB split that can hold 8-bit weights with KV room. BF16 at 52 GB is 128 GB unified or multi-GPU. Do not buy a second 3090 “for A4B 4-bit” — one 24 GB card already covers Unsloth’s 4-bit band. Buy the second card for 8-bit, for a second model, or for dual-3090 27B 200k.
Hybrid attention (sliding window + global, last layer global) is on the family README. It is why Unsloth can quote 256K on paper without 31B-dense KV. LocalRig still has no filled-256k A4B packet on 16 GB. Cap -c until you measure.
MTP (Jun 9, 2026) and QAT (Jun 5) are Unsloth add-ons. MTP is a speed claim in Studio, not a 16 GB tok/s. QAT is a different file. NVIDIA NVFP4 builds are a Blackwell/Spark stack — do not mix with this GGUF table on Ampere.
If you only have 24 GB and you wanted “the Gemma one,” also read 31B dense (17–20 GB 4-bit) and Qwen 3.8 27B. A4B is the MoE speed/quality trade Unsloth describes; it is not automatically the LocalRig homepage model.
Used RTX 3090 buying guide if you are actually purchasing the 24 GB card. Two 3090s vs one 4090 if 8-bit A4B is the real constraint.
A 16 GB card that already runs gpt-oss-20b (OpenAI 16 GB MXFP4) is not automatically a Gemma 4 26B-A4B 4-bit machine. Unsloth’s A4B 4-bit band is 16–18 GB — the same number as OpenAI’s 16 GB claim, but a different file. gpt-oss GGUFs sit ~12 GB; A4B 4-bit wants the whole 16–18 GB band plus KV. If you only have 16 GB, gpt-oss-20b is the easier guest. A4B is the maybe.
Laptop 32 GB unified (Mac / Strix Halo class) can hold Unsloth’s 4-bit A4B weights with more KV room than a 16 GB dGPU. Speed is still a shortfall: no LocalRig MLX or llama.cpp Metal packet for this SKU. Strix Halo vs Mac Studio if that is the purchase, knowing 8-bit A4B (28–30 GB) wants more than 32 GB.
Do not use the generic dense VRAM calculator as if A4B were 26B dense. Active parameters are 4B; the file is still Unsloth’s 16–18 GB 4-bit row. Calculator estimates without that table are discovery only.
Power: a 16 GB 4060 Ti is a 160 W-class part; a used 3090 is 350 W. If the only model you will run is A4B Q4, the 16 GB card is the cheaper watt. If you also want 31B dense or Qwen 3.8 27B, buy the 3090 once.
E2B and E4B are other Unsloth Gemma 4 rows (4–6 GB 4-bit). They are laptop SKUs. Landing on this URL with those files is the same SKU mix as grabbing 31B dense. Hub ~10.7M downloads on the A4B card (2026-09-10) do not make 8-bit fit one 3090.
Who this is NOT for
- People who downloaded 31B GGUFs. Dense 17–20 GB 4-bit, not this MoE.
- 8-bit shoppers on one 3090. 28–30 GB ≠ 24 GB.
- E2B/E4B laptop buyers. Different Unsloth rows (4–6 GB 4-bit).
- Anyone needing LocalRig tok/s or filled-256k on 16 GB. Shortfalls stay visible.
- Shoppers who treat 10.7M Hub downloads as a 3090 speed claim. Downloads ≠ tok/s.
Methodology
- Fit numbers: Unsloth Gemma 4 table, 2026-09-10. Vendor bands, not traces.
- SKU identity: Unsloth GGUF
base_model: google/gemma-4-26B-A4B-it, Hub MCP 2026-09-10. - Speed: none first-party.
- LocalRig first-party: none as of 2026-09-10.
Sources
Frequently Asked Questions
Can I run Gemma 4 26B-A4B on 16 GB?
Unsloth's 4-bit band is 16–18 GB total memory. A 16 GB card is at the floor for weights; KV will be tight. A 24 GB 3090 is the more comfortable used-card cell.
Is this smaller than Gemma 4 31B?
Total parameters are lower and only 4B are active, so Unsloth lists a slightly lower 4-bit band (16–18 vs 17–20 GB). It is a different architecture, not a quantized 31B.
Does 8-bit fit a 3090?
No. Unsloth's 8-bit band is 28–30 GB. One 24 GB card misses that. Dual 3090 (48 GB) is the 8-bit conversation.
How fast is it vs 31B?
Unsloth says A4B is faster because of MoE. LocalRig has no first-party tok/s on either SKU. Do not invent a 3090 ratio.
256K context on 16 GB?
Unmeasured here. File-fit is not filled-context. Leave that shortfall visible.