What Can I Run?

Can I Run Kimi K2.7-Code Locally? Still 1T / ~595 GB Q8, Not a 3090

Skip consumer GPUs; Unsloth Q8 is 595 GB — same class as K2.6
Top Pick Skip consumer GPUs; Unsloth Q8 is 595 GB — same class as K2.6

Short answer: Kimi K2.7-Code is moonshotai/Kimi-K2.7-Code: 1T / 32B active, 256K, multimodal, built on K2.6. Unsloth (GGUF README, 2026-09-10): lossless UD-Q8_K_XL is 595 GB, ~10 GB larger than UD-Q4_K_XL. That is not a 24 GB or 128 GB download. Same skip/rent class as K2.6. LocalRig has not measured tok/s.

Code SKU, same MoE floor

Moonshot’s own table matches K2.6’s shape. Official summary on the Unsloth GGUF README: 1T total, 32B active, 61 layers, 384 experts (8 per token), MLA, 256K, MoonViT 400M. Vendor pitch: long-horizon coding, ~30% fewer thinking tokens vs K2.6. Those are quality / efficiency claims. They do not turn 595 GB into 48 GB.

Hub (2026-09-10): moonshotai/Kimi-K2.7-Code ~1.027T parameters, architecture kimi_k25 — same order as K2.6. Downloads ~1.8M. Volume is not a 3090 fit.

Unsloth’s K3 docs mention K2.6 and K2.7 together when explaining lossless Q8 (MXFP4 MoE + BF16 rest). Treat K2.7-Code as K2.6-class memory, not K3’s 594 GB 1-bit floor (that 594 GB is K3 1-bit, coincidentally near K2.7 Q8 — do not mix the SKUs).

595 GB Q8 / ~585 GB Q4 vs your silicon

24 GB ≈ 4% of Q8. 128 GB ≈ 22% of Q8. Dual 3090 48 GB ≈ 8%. Unsloth’s working rule on K2.6 was RAM+VRAM ≈ quant size. K2.7-Code’s Q8 is 595 GB. There is no LocalRig packet that makes RAM-offload on a 128 GB Mac a daily driver. If you already have a 512 GB+ unified box, you are in the class Unsloth documents for K2.6 Dynamic 2-bit (350 GB) — still read the K2.7 GGUF you actually download; do not assume K2.6 340 GB 2-bit files.

Browse used RTX 3090 on eBay → for 27B, not K2.7.

Mac Studio 512GB on Amazon → only if you are actually in that memory class. 128 GB is the Flash-0731 class, not this.

Rent: RunPod, Vast.ai, GPUMart, Vultr GPU Cloud.

Runtime

llama.cpp GGUF + mmproj-BF16 (953.9 MB on the Unsloth repo). Harmony-style Kimi template still applies. Ollama: no LocalRig tag+hardware+tok/s packet. Official Hub inference providers (Novita, Fireworks, etc.) are API paths, not a 3090 recipe.

Related: Can I run Kimi K2.6 locally?. Can I run Kimi K3 locally?. Can I run Qwen 3.8 locally?.

K2.7-Code’s vendor benches (Kimi Code Bench v2 62.0 vs K2.6 50.9) are Moonshot quality tables. They do not shrink Unsloth Q8 from 595 GB to a 24 GB file. Token-efficiency (~30% fewer thinking tokens vs K2.6) helps latency and cost on a machine that already fits, not VRAM fit.

UD-IQ1_M through UD-Q4_K_XL directories exist on the Unsloth repo. This page quotes the README’s Q8 595 GB / Q4 ~585 GB pair because that is the lossless/near-lossless story they highlight. If you download UD-IQ1_M, measure that directory’s size before claiming a new floor. Do not assume 1-bit K2.7-Code equals K3’s 594 GB 1-bit — different models.

256K context on the card is extra KV on top of 595 GB. A 512 GB unified machine that “just fits” Q4 still will not hold a filled 256k at full KV precision. Shortfall.

Modified MIT on the official license badge is not Apache-2.0. Read the license if you ship a product. License is not VRAM.

K2.7-Code vs gpt-oss-20b: both say “coding,” one is 595 GB Q8 and one is 16 GB MXFP4. If the machine is a used 3090, gpt-oss-20b and Qwen 3.8 27B are the coding locals. K2.7-Code is the API-or-rent coding local.

mmproj-BF16 at 953.9 MB is noise next to 595 GB. Vision does not change the skip.

Kimi Code CLI in the eval footnotes uses 262,144 context and thinking on. That is a bench harness, like Ornith’s 48 GB RAM Harbor note — not your 3090. Production API providers on the Hub card are how most people will actually use K2.7-Code.

If you already rented a multi-GPU node for K2.6, K2.7-Code is a same-class swap, not a downsize. Confirm Unsloth Q4 vs Q8 disk before you pull.

Affiliate path for this SKU is rental, not eBay 3090. Amazon Mac 128 is the wrong K2.7 cart. 512 GB unified is the first Mac class worth discussing, and even then you size the GGUF, not the adjective “Studio.”

Hub MCP (2026-09-10): moonshotai/Kimi-K2.7-Code ~1.027T (1026879.4M), architecture kimi_k25, ~1.8M downloads. Volume and likes are not a 3090 fit. Live inference providers on that card (Novita, Fireworks, and others) are how most people will actually use a 595 GB coding MoE.

What to run instead on hardware you already own

16–24 GB: gpt-oss-20b (OpenAI 16 GB MXFP4, Harmony required) or Qwen 3.8 27B Q4. Both are coding-capable locals on a used 3090. K2.7-Code is not a “Code” shrink of K2.6; it is the same 1T planet with a coding eval table.

128 GB unified: still skip this SKU. Flash-0731 (~103 GB 3-bit) and GLM-5.3-Flash 3-bit are the 128 GB DeepSeek/GLM lanes. K2.7 Q8 at 595 GB is about 4.6× a Spark cube.

Already on K2.6 hardware (350 GB+ / 512 GB unified / multi-GPU): treat K2.7-Code as a same-class swap. Pull Unsloth Q4 vs Q8 explicitly. Do not assume K2.6’s 340 GB UD-Q2_K_XL files. Do not transplant K2.6’s Unsloth B200 >40 tok/s onto K2.7 without a new packet.

Rent if you need Moonshot’s coding table and do not already have that memory class: RunPod, Vast.ai, GPUMart, Vultr GPU Cloud. Cost framing: H100 rental price comparison.

K2.7-Code vs Kimi K3: K3’s Unsloth 1-bit floor is 594 GB on a 2.8T model. Coincidence that K2.7 Q8 is 595 GB. Different SKUs, different skip math.

Who this is NOT for

  • Shoppers who thought “Code” meant 20B. Still 1T.
  • K2.6 readers who expected a smaller GGUF. Unsloth Q8 595 GB is K2.6’s Q8 neighborhood (K2.6 Q8 disk 595 GB on that guide).
  • K3 shoppers. 2.8T / 594 GB 1-bit is a bigger skip.
  • 24 GB / 128 GB buyers.
  • Anyone needing LocalRig tok/s.

Methodology

  • Fit numbers: Unsloth K2.7-Code GGUF README Q8 595 GB / Q4 ~10 GB smaller, 2026-09-10.
  • SKU identity: Official table 1T / 32B / 256K; Hub ~1.027T.
  • Speed: none first-party.
  • LocalRig first-party: none as of 2026-09-10.

Sources

Frequently Asked Questions

Does K2.7-Code fit a 4090 or 128 GB Mac?

No. Unsloth's lossless Q8 GGUF is 595 GB, about 10 GB larger than their Q4. That is the same 350–600 GB planet as K2.6. A 24 GB or 128 GB machine cannot hold the experts.

Is K2.7-Code smaller than K2.6 because it is 'Code'?

No. The official summary is still 1T total / 32B active, 256K, MoonViT. Coding focus does not shrink the MoE.

Is this Kimi K3?

No. K3 is 2.8T / 104B active with a 594 GB 1-bit Unsloth floor. Different page.

How fast is K2.7-Code locally?

LocalRig has no first-party row. Do not transplant K2.6's Unsloth B200 >40 tok/s onto K2.7 without a new packet.

What should I run on 24 GB instead?

Qwen 3.8 27B Q4, or another 16–24 GB SKU. Not K2.7-Code.

Sources

  • unsloth/Kimi-K2.7-Code-GGUF README, https://huggingface.co/unsloth/Kimi-K2.7-Code-GGUF, Hub MCP 2026-09-10
  • Unsloth guide linked from GGUF: https://unsloth.ai/docs/models/kimi-k2.7-code
  • moonshotai/Kimi-K2.7-Code, https://huggingface.co/moonshotai/Kimi-K2.7-Code, Hub MCP 2026-09-10 (~1.027T)