Can I Run Kimi K2.6 Locally? 1T MoE, 350 GB Floor, Not a 48 GB Card
Short answer: Kimi K2.6 is moonshotai/Kimi-K2.6: 1T total / 32B active, native multimodal, 256K context. Unsloth’s Dynamic 2-bit (UD-Q2_K_XL) is ~340 GB on disk and they want 350 GB+ of RAM+VRAM. A used 3090, a 4090, dual 3090s, and a 128 GB Mac Studio all miss that floor. The honest paths are a 512 GB-class RAM/GPU box, a multi-GPU datacenter node, or rent. LocalRig has not measured tok/s on this SKU.
The previous version of this URL used a fake card (Qwen/Kimi-k2-6) and a 48–72 GB Q4 table. That math does not describe this model. If you sized a dual-3090 build from the old copy, keep the 3090s for Qwen 3.8 27B — not for K2.6.
What is actually on the Kimi K2.6 card?
A 1T MoE with 32B active, 256K context, and a vision encoder — not a 14B dense coder. Hugging Face Hub (accessed 2026-09-10) lists moonshotai/Kimi-K2.6 at ~1.027T parameters, architecture kimi_k25, pipeline image-text-to-text, ~8.1M downloads. Moonshot’s README table: 61 layers, 384 experts (8 selected per token), MLA attention, 256K context, MoonViT 400M. License: Modified MIT.
Vendor positioning (Moonshot, not a LocalRig ranking): long-horizon coding, coding-driven design, agent swarm. Those claims do not change the VRAM floor. All experts still have to live in memory.
Do not confuse K2.6 with later Kimi SKUs (K2.7-Code, K3). Those get their own skip/rent pages. This URL is K2.6 only.
How much memory does Unsloth actually list?
350 GB-class at Dynamic 2-bit; ~600 GB at Q4/Q8. Copied from Unsloth (accessed 2026-09-10). Units are total memory (RAM + VRAM, or unified). Disk sizes:
| Measurement | Dynamic 2-bit (UD-Q2_K_XL) | Q4 (UD-Q4_K_XL) | Q8 (UD-Q8_K_XL, lossless) |
|---|---|---|---|
| Disk space | 340 GB | 584 GB | 595 GB |
| Working rule they state | 350 GB+ RAM/VRAM | ~600 GB class | ~600 GB class |
Unsloth’s quality note: Kimi stores MoE weights as int4 and everything else in BF16; UD-Q8_K_XL follows that layout and they call it lossless. UD-Q4_K_XL is near full precision and still a ~600 GB problem. Their published perplexity on that page (2.4131 at 2-bit vs ~1.842 at Q4/Q8) is a quantization-quality table, not a hardware tok/s table.
Suggested context in Unsloth’s usage guide: 98,304 (up to 262,144). Thinking mode: temperature 1.0, top_p 0.95. Instant/non-thinking: temperature 0.6, top_p 0.95. GGUFs support vision via mmproj-F16. None of that fits in 24 GB.
Can a 24 GB, 48 GB, or 128 GB machine run Kimi K2.6?
No, not at Unsloth’s documented quants.
- 24 GB (3090/4090): about 7% of the 350 GB 2-bit floor. Skip.
- Dual 3090 (48 GB): still about 14% of 350 GB. Skip.
- 128 GB Mac / Spark: the right class for DeepSeek-V4-Flash-0731 3-bit (~103 GB). For K2.6, 128 GB is still under half of Dynamic 2-bit. Offload-to-RAM will crawl; LocalRig has no packet that makes that a daily driver.
Browse used RTX 3090 24GB on eBay → if you are buying for 27B. Do not buy a 3090 for K2.6.
Check Mac Studio listings on Amazon → only if you are actually shopping a 512 GB-class unified-memory machine. A 128 GB SKU is the wrong K2.6 buy.
What is the honest local or cloud path?
Rent a multi-GPU datacenter node, or own a 350–600 GB memory pool. Unsloth’s speed sentence is: if it fits, >40 tok/s on B200s. Community-cited, B200-only. No LocalRig Spark/Mac/3090 figure.
llama.cpp path they document: latest llama.cpp, -hf unsloth/Kimi-K2.6-GGUF:UD-Q2_K_XL, plus mmproj-F16 for vision. That command still needs the 350 GB. Official serving (Novita, Fireworks, etc. on the Hub card) is an API path, not a 24 GB recipe.
Try-before-buy on large GPUs: RunPod, Vast.ai, GPUMart GPU hosting. Datacenter catalog: Vultr GPU Cloud. Break-even math: rent vs buy.
If the job is daily local coding on hardware you already own, stay on Qwen 3.8 27B Q4 or GLM-5.3-Flash only if you have the 128 GB class that Flash actually fits. K2.6 is the wrong download for those machines.
What about llama.cpp, Ollama, and “it fits if I offload”?
Offload is not a 350 GB skip. llama.cpp will load a GGUF that does not fit in VRAM by paging into system RAM (and then disk). Unsloth says that still works, “just slower due to offloading.” LocalRig has no named tok/s packet for K2.6 on a 128 GB Mac with RAM offload, so this page does not treat that as a daily-driver cell. If you already own a 512 GB unified-memory machine, Unsloth’s UD-Q2_K_XL (340 GB disk, 350 GB+ working rule) is the band they document as the size/quality balance. Q4/Q8 stay a ~600 GB problem.
Ollama: LocalRig has no ollama run library tag plus hardware plus tok/s packet for K2.6 as of 2026-09-10. A blog one-liner without those fields is not a cell. The GGUF path Unsloth documents is llama.cpp with unsloth/Kimi-K2.6-GGUF:UD-Q2_K_XL and mmproj-F16 for vision. Thinking vs instant sampling differs (1.0 vs 0.6 temperature). That is vendor procedure, not a LocalRig lab notebook.
Vision is on the official card (MoonViT 400M) and on Unsloth’s GGUF via mmproj. A text-only coding loop still has to hold the 1T experts. Do not budget “text-only therefore 48 GB.”
Hub downloads (~8.1M on the official card, ~883K on the Unsloth GGUF, accessed 2026-09-10) are volume, not a ranking against DeepSeek-V4-Flash-0731 or Qwen 3.8 27B. K2.6 is the larger memory class. Flash-0731 is the 128 GB-class DeepSeek SKU. They are not peers on a used 3090.
Who this is NOT for
- Anyone using the old 48–72 GB table on this URL. That copy cited a non-existent Qwen repo and understated the file by ~6×.
- 24 GB and 48 GB shoppers. Wrong SKU. Buy or keep a 3090 for 27B, not for K2.6.
- 128 GB Mac / Spark buyers who think “unified memory” means 1T MoE fits. 128 GB ≠ 350 GB.
- Readers who want LocalRig tok/s on K2.6. There is no first-party row. Unsloth’s >40 tok/s figure is B200.
- People collapsing K2.6 into K2.7-Code or K3. Different SKUs; do not reuse this math.
Methodology
- Fit numbers: Unsloth Kimi K2.6 guide and GGUF repo, accessed 2026-09-10. Vendor disk sizes and 350 GB working rule. Not LocalRig measurements.
- SKU identity:
moonshotai/Kimi-K2.6via Hugging Face Hub MCP (2026-09-10): ~1.027T parameters, 32B active in the official README table, 256K context, MoonViT. - Speed numbers: Unsloth “>40 tok/s on B200s if it fits.” Not transferable to 3090 / Spark / Mac.
- LocalRig first-party: none for Kimi K2.6 as of 2026-09-10.
Sources
- moonshotai/Kimi-K2.6 — official card, 1T / 32B, 256K, multimodal (Hub MCP 2026-09-10).
- Unsloth Kimi K2.6 local guide — 340/584/595 GB disk, 350 GB+ working rule, B200 >40 tok/s, sampling (accessed 2026-09-10).
- unsloth/Kimi-K2.6-GGUF — Dynamic GGUF + mmproj.
- Can I run DeepSeek-V4-Flash-0731 locally? — 128 GB-class Flash, not 1T.
- Can I run Qwen 3.8 locally? — 24 GB dense daily driver.
- H100 rental price comparison — datacenter rental, not a 3090 cell.
Frequently Asked Questions
Does Kimi K2.6 fit on a single RTX 4090?
No. Unsloth's smallest practical Dynamic 2-bit GGUF is about 340 GB on disk and they want 350 GB+ of RAM+VRAM. A 24 GB 4090 cannot hold the experts, active or not.
Will a 128 GB Mac Studio run Kimi K2.6?
Not at Unsloth's UD-Q2_K_XL size. 128 GB is about one-third of 350 GB. A 128 GB Mac is the right class for DeepSeek-V4-Flash-0731 3-bit, not for K2.6.
Is the Hugging Face repo Qwen/Kimi-k2-6?
No. The official card is moonshotai/Kimi-K2.6. The older LocalRig source line pointing at Qwen/Kimi-k2-6 was wrong.
How fast is Kimi K2.6 locally?
LocalRig has no first-party row. Unsloth says if the model fits you will see >40 tok/s on B200s. That is a B200 packet, not a 3090, not a Mac. Do not transplant it.
What should I run instead on 24 GB?
Qwen 3.8 27B Q4 on a used 3090. Kimi K2.6 is the wrong download for that machine.