What Can I Run?

Can I Run gpt-oss-20b Locally? Native MXFP4 in 16 GB, Harmony Required

16 GB class for MXFP4 weights; used 3090 24GB for KV headroom
Top Pick 16 GB class for MXFP4 weights; used 3090 24GB for KV headroom

Short answer: openai/gpt-oss-20b is 21B total / 3.6B active, Apache-2.0, 128K context, native MXFP4 on MoE layers. OpenAI: it runs within 16 GB. Unsloth (accessed 2026-09-10): ≥14 GB for 6+ tok/s on Dynamic 4-bit. Hub ls: GGUFs ~11.5–13.8 GB. A 16 GB card is the floor. A used 3090 is headroom, not a requirement. Harmony is required. This is not gpt-oss-120b. LocalRig has not measured tok/s on a 3090.

What is gpt-oss-20b?

OpenAI’s smaller open-weight MoE, not a dense 20B Llama clone. Hub (2026-09-10): ~20.91B parameters, architecture gpt_oss, ~96M downloads. Card highlights: configurable reasoning effort, full chain-of-thought, Apache-2.0, MXFP4 so 20b fits 16 GB and 120b fits one H100. Use Harmony (openai/harmony). Unsloth’s F32 GGUF note: MXFP4 upcast to BF16 per layer for the unquantized file.

Vendor “rivals o3-mini” language is a quality claim, not a VRAM table.

GGUF sizes — why Q4 and Q8 look similar

Experts are already 4-bit-class. Hub ls of unsloth/gpt-oss-20b-GGUF (2026-09-10):

QuantFile size
Q4_K_M11.6 GB
UD-Q4_K_XL11.9 GB
Q8_012.1 GB
UD-Q8_K_XL13.2 GB
F1613.8 GB

You do not “Q4 it down to 6 GB” like a dense 20B. Unsloth’s 14 GB / 6+ tok/s sentence is a memory-vs-speed rule of thumb, hardware-unspecified beyond unified/RAM. Community-cited, not LocalRig.

16 GB vs used 3090

16 GB matches OpenAI’s 16 GB claim for weights. 24 GB is KV and less offload. A 4060 Ti 16 GB is on-class for this SKU at short context. A 3090 still makes sense if you already own one or want longer windows — it is not a 350 GB Kimi buy.

Browse used RTX 3090 24GB on eBay →

Check RTX 4060 Ti 16GB on Amazon →

Rent: RunPod, Vast.ai, GPUMart.

Ollama vs llama.cpp vs vLLM

All three are named on the official card. Harmony still applies. Ollama: ollama pull gpt-oss:20b / ollama run gpt-oss:20b (OpenAI cookbook path). Unsloth llama.cpp: -hf unsloth/gpt-oss-20b-GGUF:F16 in their tutorial. vLLM: OpenAI documents a gptoss wheel index — pin versions; that is serving, not a 3090 tok/s. LocalRig has no filled-128k packet.

Related: Ollama vs llama.cpp vs vLLM. How to run LLMs locally.

16 GB is the story — 24 GB is comfort, 120b is a different planet

OpenAI sized 20b for 16 GB on purpose. The GGUF files sitting at ~12 GB are a consequence of native MXFP4, not a LocalRig estimate. A used 3090 still helps: longer 128K-class windows, less offload, fewer “6+ tok/s only if you have 14 GB free” surprises. Unsloth’s 6+ tok/s / 14 GB line is vendor, hardware-unspecified. Do not paste it onto a 3060 12 GB card and call it a cell — 12 GB is under both OpenAI’s 16 GB and Unsloth’s 14 GB rule of thumb once the OS is running.

gpt-oss-120b is the sibling Unsloth documents at ≥66 GB for 1-bit / 6+ tok/s. One 3090 cannot hold 120b. This page does not size 120b. Harmony still applies on both.

Ollama’s gpt-oss:20b tag is the lowest-friction path OpenAI names for consumer hardware. llama.cpp via Unsloth F16 GGUF (13.8 GB) is the file-level path. vLLM’s special wheel is a server path with version pins — treat it as datacenter-or-lab, not a 3090 default. LocalRig has no Ollama-vs-llama.cpp tok/s triangle for this SKU on Ampere.

Reasoning effort (low/medium/high) changes latency and KV, not the 11.6 GB Q4_K_M file. Full chain-of-thought is on the card as a debugging feature, not something to show end users — and every CoT token is still cache.

If you already own a 3090 for Qwen 3.8 27B, gpt-oss-20b is an easier VRAM guest on that card. If you are buying a new 16 GB card only for 20b, that matches OpenAI’s 16 GB claim better than buying a 3090 “because LocalRig likes 3090s.” The 3090 remains the used-market workhorse when you also want 27B.

Used RTX 3090 buying guide. Best GPU under $500 if 16 GB used is the budget.

Harmony is the load-bearing constraint people skip. If a frontend sends ChatML or a raw Llama template, the model will look broken. Ollama and Unsloth llama.cpp paths apply Harmony when you use their gpt-oss tags. Rolling your own server means you own that template. LocalRig is not documenting a custom template here.

12 GB cards (3060 12 GB, 4070 12 GB): under OpenAI’s 16 GB and under Unsloth’s 14 GB 6+ tok/s rule of thumb. Not a LocalRig cell. 8 GB is a skip.

gpt-oss-20b vs Ornith 9B: Ornith Q4 is 5.7 GB dense; gpt-oss is ~12 GB MXFP4 MoE. Both fit 16 GB; they are not the same quality class and LocalRig has tok/s for neither on Ampere.

Apache-2.0 is why this SKU shows up in commercial local stacks. License is not VRAM. 128K context on the card is the max; filled 128k on 16 GB is a shortfall.

Function calling, browser, and Python tools on the OpenAI card are capabilities, not extra weight files, but tool traces still land in KV. A 16 GB card running high-reasoning + tools will feel the cache before a 3090 does. Keep effort on low if you are VRAM-bound.

Colab notebooks Unsloth advertises for 20b fine-tunes are cloud T4/A100 paths, not a 3090 inference cell.

Who this is NOT for

  • gpt-oss-120b shoppers. Different Unsloth floor (~66 GB 1-bit).
  • People skipping Harmony. Official: it will not work correctly.
  • Shoppers expecting dense-20B Q4 file sizes. MXFP4 already compressed the experts; files sit ~12 GB.
  • Anyone needing LocalRig 3090 tok/s. Unsloth’s 6+ tok/s / 14 GB is not a 3090 llama-bench.
  • Kimi/DeepSeek Flash buyers. Those are 100 GB+ MoEs. This SKU is the 16 GB class.

Methodology

  • Fit numbers: OpenAI 16 GB MXFP4 claim; Unsloth ≥14 GB / 6+ tok/s; Hub ls file sizes 2026-09-10.
  • SKU identity: openai/gpt-oss-20b Hub MCP (~20.91B, 3.6B active on the card).
  • Speed: Unsloth 6+ tok/s at ≥14 GB labeled vendor; no LocalRig 3090 row.
  • LocalRig first-party: none as of 2026-09-10.

Sources

Frequently Asked Questions

Can I run gpt-oss-20b on 16 GB?

OpenAI's card says native MXFP4 MoE lets gpt-oss-20b run within 16 GB. Unsloth wants at least 14 GB unified/RAM for 6+ tok/s on their Dynamic 4-bit. A 16 GB card is the documented floor for weights. LocalRig has no first-party tok/s.

Do I need a 3090?

Not for the weight file. Unsloth GGUFs cluster around 11.5–13.8 GB because experts stay MXFP4. A 24 GB 3090 is extra KV and easier offload, not a 70B-class buy.

Is this gpt-oss-120b?

No. 120b is the H100-class sibling (Unsloth 1-bit ~66 GB). This page is 20b only.

Can I skip Harmony?

OpenAI says both models were trained on Harmony and will not work correctly otherwise. Use a runtime that applies the template (Ollama, llama.cpp with Unsloth flags, Transformers chat template).

Ollama one-liner?

OpenAI's card documents `ollama pull gpt-oss:20b` then `ollama run gpt-oss:20b`. That is a named runtime path. LocalRig still has no first-party tok/s on a 3090.

Sources

  • openai/gpt-oss-20b card, https://huggingface.co/openai/gpt-oss-20b, Hub MCP 2026-09-10
  • Unsloth gpt-oss guide, https://docs.unsloth.ai/models/gpt-oss-how-to-run-and-fine-tune, accessed 2026-09-10
  • unsloth/gpt-oss-20b-GGUF file sizes via Hub MCP, 2026-09-10