CAN I RUN IT · SEPTEMBER 2026 INDEX
Pick the memory class, then the SKU.
This cluster is a constraint library, not a harvest page. Qwen 3.8 27B Q4 stays the 24 GB dense default. Flash-Next, GLM-5.3-Flash, and DeepSeek-V4-Flash-0731 are 128 GB-class. Kimi and DeepSeek Pro are skip/rent. Empty tok/s cells are left empty.
24 GB class
Used RTX 3090 / 16–24 GB. Qwen 3.8 27B Q4 is the dense daily driver. Other SKUs on this row are guests on the same card, not replacements for that default.
- Can I Run Qwen 3.8 Locally? 27B vs Flash-Next vs the Giant MoE Qwen 3.8 is three different local SKUs. This guide maps Qwen3.8-27B dense to 24 GB cards, Flash-Next to Spark and large unified memory, and the 125B / 2.4T MoEs to machines you probably should not buy for this model. Requirement math first; community tok/s second.
- Qwen 3.8 27B Q4 on a Used RTX 3090: The Short-Session Path A used RTX 3090 still runs Qwen3.8-27B Q4 at 4k–32k. This page is the named 24 GB cell: Unsloth weight bands, llama.cpp fit limits, and community tok/s with URL, date, quant, and runtime. It is not the 200k agent default.
- Qwen 3.8 27B Q4 at 200k: Spark, Mac 128, or Dual 3090? At a 200k agent window, Qwen3.8-27B Q4 is a KV-cache decision, not a weights decision. This page maps Spark 128 GB, Mac 128 GB, and dual RTX 3090 against named public cite packets. A single 24 GB 3090 is the short-session path.
- Can I Run GLM-4.7-Flash Locally? 30B-A3B on 24 GB, Not GLM-5.3 GLM-4.7-Flash is Z.ai's 30B-A3B MoE, not GLM-5.3-Flash and not GLM-5.2. Unsloth's 4-bit path wants ~18 GB; they say it runs on 24 GB. UD-Q4_K_XL is 17.5 GB. A used 3090 is the cell; 5.3-Flash is 100 GB-class.
- Can I Run Gemma 4 31B Locally? Unsloth 4-bit 17–20 GB on a 24 GB Card google/gemma-4-31B-it is a ~31B dense multimodal SKU with 256K context, not the 26B-A4B MoE. Unsloth wants 17–20 GB total memory at 4-bit. A used 3090 is the 24 GB cell; 16 GB is offload-or-skip at that band.
- Can I Run Gemma 4 26B-A4B Locally? 16–18 GB 4-bit MoE vs 31B Dense Gemma 4 26B-A4B is a 256K MoE (4B active), not the 31B dense SKU. Unsloth's 4-bit band is 16–18 GB. That is a 16 GB maybe and a 24 GB used-3090 yes for weights; 8-bit wants 28–30 GB.
- Can I Run gpt-oss-20b Locally? Native MXFP4 in 16 GB, Harmony Required openai/gpt-oss-20b is 21B / 3.6B active with native MXFP4. OpenAI says it runs within 16 GB. Unsloth GGUFs are ~11.5–13.8 GB. A 16 GB card is the floor; a used 3090 is KV headroom. Harmony format is required.
- Can I Run Ornith-1.0-9B Locally? 9B Dense on 16 GB vs a Used 3090 Ornith-1.0-9B is a dense ~9B coding SKU, not a 100 GB MoE. Unsloth's Q4_K_M GGUF is 5.7 GB; BF16 is 17.9 GB. 16 GB and used 3090 both hold Q4; 24 GB is the longer-context card. Not Ornith-35B.
- Can I Run Ornith-1.0-35B Locally? Q4 ~22 GB on 24 GB, Q8 Wants Dual 3090 Ornith-1.0-35B is the MoE sibling of Ornith 9B, not a 9B download. Unsloth UD-Q4_K_XL is 22.3 GB; Q8_0 is 36.9 GB. A used 3090 is a tight Q4 cell; dual 3090 is the Q8 conversation. Hub parameter overview is a shortfall.
- Can I Run Qwen3.6-27B Locally? Prior 24 GB Driver, Not Qwen 3.8 Qwen/Qwen3.6-27B is the prior 27B multimodal daily driver, not Qwen 3.8. Unsloth's 4-bit band is 18 GB; 6-bit 24 GB. A used 3090 is the Q4 cell. MTP ~160 tok/s is Unsloth on RTX 6000, not a 3090 number.
128 GB unified
Mac / Spark class. Unsloth 1-bit and 3-bit bands, not a used 3090. Empty Spark/Mac tok/s cells stay empty.
- Can I Run Qwen 3.8 Flash Locally? 75 GB GGUF, Not the 27B Homepage SKU Qwen 3.8 Flash is the 125B MoE / large-RAM path, not Qwen3.8-27B on a 24 GB card. Unsloth's Flash-Next GGUF is 75 GB at 1-bit and they ask for a 96 GB machine. Public speed cites are RTX PRO 6000 MTP and a 4090-plus-RAM offload run — not LocalRig benches.
- Can I Run GLM-5.3-Flash Locally? 320B-A18B Memory Math, Not a 24 GB Card GLM-5.3-Flash (ox-alpha) is a 320B-A18B MoE, not GLM-5.2 and not a 3090 download. Unsloth's table puts 1-bit around 100 GB total memory and 3-bit on 128 GB Mac / Spark setups. Public llama-bench figures are 1× B200. Requirement math first; named benches second.
- Can I Run DeepSeek-V4-Flash-0731 Locally? 284B MoE, Not a 14B Card DeepSeek-V4-Flash-0731 is a 284B-class MoE with 13B active, not a 14B dense download. Unsloth's 3-bit GGUF is ~103 GB and wants ~110 GB total memory. Spark or Mac 128 at 3-bit, or rent; a used 3090 is the wrong buy.
- Can I Run MiniMax M3 Locally? 428B MoE, Not an 8–13B KV Essay MiniMax-M3 is a ~428B / 23B-active multimodal MoE with 1M context, not an 8–13B model. Unsloth's 1-bit GGUF wants ~133 GB; 3-bit 164–200 GB. Spark or Mac 128 is a tight 1-bit maybe; a 3090 is a skip. GGUF is experimental.
Skip / rent
Hundreds of gigabytes of weights. Do not size a 3090. Rent or skip; the 24 GB and 128 GB SKUs above are the local path.
- Can I Run Kimi K2.6 Locally? 1T MoE, 350 GB Floor, Not a 48 GB Card moonshotai/Kimi-K2.6 is a 1T-parameter multimodal MoE (32B active), not a 48–72 GB 14B peer. Unsloth's Dynamic 2-bit GGUF wants ≥350 GB. Skip 24 GB and 128 GB boxes; rent or a 512 GB-class machine.
- Can I Run Kimi K2.7-Code Locally? Still 1T / ~595 GB Q8, Not a 3090 moonshotai/Kimi-K2.7-Code is a 1T / 32B-active coding SKU built on K2.6. Unsloth lossless Q8 is 595 GB; Q4 is ~10 GB smaller. Skip 24 GB and 128 GB. Same memory planet as K2.6. Rent, do not size a 3090.
- Can I Run Kimi K3 Locally? 2.8T / 594 GB 1-bit, Not K2.6 moonshotai/Kimi-K3 is 2.8T / 104B active with 1M context. Unsloth Dynamic 1-bit is 594 GB and they want ~610 GB total memory. Skip 24 GB, 128 GB, and K2.6's 350 GB class. Rent. Do not size a 3090.
- Can I Run DeepSeek-V4-Pro-0813 Locally? 1.57T Skip/Rent, Not Flash-0731 DeepSeek-V4-Pro-0813 is 1.57T / 48B active, not Flash-0731's ~103 GB 3-bit. Unsloth Q4 GGUF is a ~20-shard ~850 GB-class download. Skip 24 GB and 128 GB. Rent a datacenter node. Do not size a 3090.
- Can I Run GLM-5.2 Locally? Hardware Requirements for the 1M-Context Flagship GLM-5.2 (released June 2026) is open-weight but not runnable on most home hardware. The 1M-context model demands a 256GB-class machine and hits single-digit tokens/sec even then. Honest constraints, the math, and where to go instead.
Reference
Size pickers, quantization, and older family pages. Not the September SKU map.
- Which Quant Should You Download? Q4_K_M vs Q8_0 vs F16, Decided by Your Hardware The download page forces a choice — Q4_K_M, Q8_0, F16, or an IQ variant — and most guides answer it with vibes. This is a constraint selector: your VRAM/unified memory sets the ceiling, your quality tolerance sets the floor, and for almost everyone the answer is Q4_K_M.
- GGUF vs GPTQ vs AWQ: Which Model Format for Your Setup? Format confusion is runtime confusion. GGUF belongs to llama.cpp's world (CPU offload, Macs, single-GPU flexibility), GPTQ/AWQ to vLLM (batching, multi-GPU, datacenter efficiency), and EXL3 to the enthusiast edge. This glossary maps format to runtime to hardware so you download the right artifact for your stack.
- Can I Run Qwen3.5 Locally? Picking Your Size from 3B to 235B Qwen3.5 spans 3B to 235B parameters. This guide pairs your hardware to the right model size using requirement math, not marketing claims. A single table from 8GB cards to 256GB Macs, with Unsloth Dynamic 2.0 quantization recommendations.
- Hardware to Run a 32B Model Locally: The Sweet Spot Tier 32B models hit the 2026 sweet spot: near-70B quality at half the hardware cost and power draw. This guide covers the VRAM math, which cards qualify (used 3090, A6000, 4090, some Macs), and when 32B at Q5 outperforms 70B at Q2 in practical work.
- Hardware to Run a 70B Model Locally: VRAM, the 48GB Wall, and Your Real Options What it actually takes to run a 70B model at home: the VRAM math, why 48GB is the practical floor at Q4, and the five hardware paths (dual 3090, used A6000, Apple Silicon, DGX Spark, or cloud).
- Hardware to Run a 7B/8B Model Locally: RTX 3090, Apple M3 Max, and Budget Options Benchmark-backed hardware guide for running 7B and 8B parameter models locally. Covers RTX 3090, Apple M3 Max, RTX 3060, and Apple M4 — with first-party Apple M4 benchmarks, community throughput data, VRAM requirements, and honest trade-offs.
- The Local-AI Hardware Buying Framework A constraint-first framework for choosing hardware to run AI models locally. Covers VRAM, memory bandwidth, quantization, Apple Silicon, DGX Spark, and budget paths — so you buy once and regret nothing.
- Quantization: What It Means for Local AI and Why It Matters Quantization reduces the numerical precision of a model's weights to shrink its memory footprint — the single technique that determines whether a 7B or 70B model fits in your GPU's VRAM and how fast it will run.