Can I Run MiniMax M3 Locally? 428B MoE, Not an 8–13B KV Essay
Short answer: MiniMax-M3 is MiniMaxAI/MiniMax-M3: ~428B total / ~23B active, native text / image / video, 1M context, MoE + MiniMax Sparse Attention. Unsloth’s 1-bit GGUF is 128 GB on disk and they want ~133 GB total memory. A used 3090 cannot load it. A 128 GB Mac / Spark is a tight 1-bit maybe, not a 3-bit daily driver. The GGUF path is experimental (llama.cpp PR 24523; MSA not supported). LocalRig has not measured tok/s.
This page replaces the earlier essay that treated M3 as an 8–13B model whose 1M window was “just KV cache.” Weights for this SKU are already in the 100 GB+ class before you talk about a million tokens.
What is MiniMax-M3 on the official card?
A ~428B multimodal MoE, not a small dense coder. Hugging Face Hub (accessed 2026-09-10) lists ~427.0B parameters, architecture minimax_m3_vl, pipeline image-text-to-text, ~636K downloads. Official README: ~428B total, ~23B active, 128 experts (4 active per token), 60 layers, 1M context, MiniMax Community License, arXiv 2606.13392.
Vendor highlights (MiniMax, not a LocalRig ranking): mixed-modality training from step one; MiniMax Sparse Attention for long context (they claim 9× prefill / 15× decode vs M2 at 1M vs GQA); coding and cowork. Unsloth repeats vendor scores (SWE-Bench Pro 59%, Terminal-Bench 2.1 66%) — those are quality figures, not a 24 GB fit table.
Unsloth also states bf16 weights ~855 GB. That is the unquantized floor, not a GGUF.
Do not mix MiniMax-M3 with MiniMax-M2.7 or other MiniMax SKUs. This URL is M3.
How much memory does Unsloth actually list?
~133 GB at 1-bit, 164–200 GB at 3-bit, 213–270 GB at 4-bit. Copied from Unsloth (accessed 2026-09-10). Units are total memory: RAM + VRAM, or unified. File size excludes KV.
| 1-bit | 2-bit | 3-bit | 4-bit | 5-bit | 8-bit |
|---|---|---|---|---|---|
| 133 GB | 148 GB | 164–200 GB | 213–270 GB | 325 GB | 460–470 GB |
They name UD-IQ1_M as the smallest current quant (128 GB disk, ≥133 GB RAM). They say UD-IQ3_XXS (159 GB) is the better quality target. UD-IQ4_XS 208 GB, UD-Q4_K_XL 265 GB — 256 GB+ or multi-GPU / CPU offload.
Sampling they copy from MiniMax: temperature 1.0, top_p 0.95, top_k 40. Default system prompt names MiniMax-M3.
Can a 24 GB card run MiniMax M3?
No. A 24 GB 3090/4090 is about one-fifth of Unsloth’s 1-bit floor. Dual 3090s (48 GB) still miss 133 GB. This is not “truncate KV to 128K on an 8B model.” The 1-bit weights already overflow consumer cards.
If your silicon is 24 GB, use Qwen 3.8 27B Q4. Browse used RTX 3090 24GB on eBay → for that job, not for M3.
Does 128 GB unified actually fit?
1-bit only, and Unsloth wants 133 GB against a 128 GB device — that is a shortfall, not a promise. A 128 GB Mac Studio or DGX Spark matches the class they describe for UD-IQ1_M, with no spare for a large KV cache and no MSA on the experimental GGUF. 3-bit (164–200 GB) wants more than 128 GB unified.
LocalRig has no Mac or Spark tok/s for M3. Do not invent one.
Check Mac Studio 128GB on Amazon → if you are shopping that class knowing 1-bit is tight. DGX Spark vs dual RTX 3090 still leaves dual-3090 under the weight file.
For 3-bit / 4-bit or official vLLM/SGLang: rent. RunPod, Vast.ai, GPUMart GPU hosting. Datacenter GPUs: Vultr GPU Cloud.
What does “experimental GGUF” mean for runtime?
llama.cpp support is a specific PR; MSA is off; the GGUF is text-only today. Unsloth and the GGUF README (accessed 2026-09-10): build llama.cpp from PR #24523, not a random release tag. llama-cli -hf unsloth/MiniMax-M3-GGUF:UD-IQ1_M. MiniMax Sparse Attention is not supported yet, so attention falls back to dense. Unsloth tells you to keep --ctx-size modest because dense attention at long context burns memory. The native 1M window on the vendor card is not the GGUF path.
Official local/serving pointers on the MiniMax card: SGLang cookbook, vLLM recipes, Transformers, KTransformers, Unsloth, ATOM (ROCm). Those are named stacks, not a 3090 Ollama one-liner. LocalRig has no Ollama library packet for M3 as of 2026-09-10.
Related: Ollama vs llama.cpp vs vLLM. How to run LLMs locally.
Why the old 1M-context KV essay was the wrong constraint
On an 8–13B dense model, a million-token window can dominate VRAM. On MiniMax-M3, the weights already do. Unsloth’s 1-bit file is 128 GB. That is before KV. The vendor card’s MSA is how MiniMax claims 1M is tractable versus GQA. The experimental GGUF does not have MSA, so you do not get those 9×/15× long-context speedups on the llama.cpp PR path. Treat 1M as a vendor-card feature for SGLang/vLLM/KTransformers, and treat GGUF as a short-to-medium context experiment until MSA lands.
Official serving docs on the card name SGLang, vLLM, Transformers, KTransformers, Unsloth, and ATOM (ROCm). Those stacks still need a memory pool in the 133 GB+ class (1-bit) or 164–270 GB (3–4-bit). A single 80 GB H100 does not hold Unsloth’s 3-bit band. Rent a multi-GPU node or a 256 GB box if 4-bit is the target. Hub downloads (~636K on the official card, ~63K on the Unsloth GGUF) are volume, not a reason to size a 3090.
Thinking vs non-thinking: MiniMax documents enabled / adaptive / disabled on the official card; Unsloth copies temperature 1.0, top_p 0.95, top_k 40. That is sampling, not a VRAM table.
If you only have 24 GB and you wanted “frontier coding locally,” the current LocalRig dense path is still Qwen 3.8 27B Q4. MiniMax-M3 is the wrong download for that card, even at 4k context.
Who this is NOT for
- Anyone who kept the old 8–13B KV essay from this URL. M3’s weights are a 100 GB+ problem before 1M context.
- 24 GB and 48 GB shoppers. Wrong SKU.
- People who need MSA and native video on GGUF today. Unsloth marks the GGUF experimental, text-only, dense-attention fallback.
- Readers who want LocalRig tok/s. No first-party row. No invented Mac number.
- Shoppers collapsing M3 into MiniMax-M2.7 or into GLM-5.2. Different labs, different files. GLM-5.2 already has its own selector.
Methodology
- Fit numbers: Unsloth MiniMax M3 guide and
unsloth/MiniMax-M3-GGUF, accessed 2026-09-10. Vendor total-memory bands. Not LocalRig measurements. - SKU identity:
MiniMaxAI/MiniMax-M3via Hugging Face Hub MCP (2026-09-10): ~427B listed parameters; README ~428B / ~23B active, 1M context, MSA, multimodal. - Runtime: Unsloth + GGUF README: experimental, PR 24523, MSA unsupported, text-only GGUF. Official SGLang/vLLM links on the card — no LocalRig hardware tok/s.
- LocalRig first-party: none for MiniMax-M3 as of 2026-09-10.
Sources
- MiniMaxAI/MiniMax-M3 — official card (Hub MCP 2026-09-10).
- Unsloth MiniMax M3 local guide — 133 GB 1-bit table, 128 GB
UD-IQ1_M, experimental GGUF notes (accessed 2026-09-10). - unsloth/MiniMax-M3-GGUF — GGUF README, PR 24523, dense-attention fallback.
- llama.cpp PR 24523 — MiniMax-M3 GGUF support branch.
- Can I run DeepSeek-V4-Flash-0731 locally? — another 128 GB-class MoE, different lab.
- Can I run Qwen 3.8 locally? — 24 GB dense path.
- Rent vs buy a GPU.
Frequently Asked Questions
Can I run MiniMax M3 on a used RTX 3090?
No. Unsloth's 1-bit band is about 133 GB of total memory. A 24 GB card cannot hold the weights. This is not an 8–13B KV-cache problem.
Will a 128 GB Mac Studio or DGX Spark run MiniMax M3?
Unsloth's smallest GGUF (UD-IQ1_M) is 128 GB on disk and they want at least 133 GB. A 128 GB unified box is the local class they describe for 1-bit, and it is tight. 3-bit wants 164–200 GB. LocalRig has no first-party tok/s.
Is MiniMax M3 an 8–13B model?
No. That was a LocalRig error on this URL. The official card is ~428B total / ~23B active, multimodal, 1M context.
Does the GGUF support 1M context and video?
The vendor card does. Unsloth's GGUF is experimental, currently text-only, and MiniMax Sparse Attention is not supported yet in that llama.cpp PR — inference falls back to dense attention. Keep context modest on the GGUF path.
How fast is MiniMax M3 locally?
LocalRig has no first-party row. Do not invent Mac or 3090 tok/s. Official serving docs point at SGLang, vLLM, KTransformers, and Unsloth — pin a hardware config before quoting speed.