Qwen 3.8 27B at 200k: Ollama vs llama.cpp vs vLLM
Short answer: pick the stack, then the runtime. Ampere 3090 stays on GGUF (Ollama or llama.cpp) and will not fully GPU-reside a 200k f16-KV session on one 24 GB card. Spark / Blackwell is where vLLM (or SGLang) NVFP4 recipes reserve 262,144. The category comparison Ollama vs llama.cpp vs vLLM still routes by concurrent users. This page is Qwen3.8-27B × 200k.
LocalRig has no 27B runtime bake-off. Every tok/s is community-cited. The 8B M4 first-party row stays on the category page.
What each runtime can actually claim at this window
| Runtime | Model format for 27B | 200k / 262k packet in hand | Hardware the packet used | Honest use |
|---|---|---|---|---|
| Ollama | Library GGUF qwen3.8:27b 18 GB, advertised 256K (library) | Advertised window only. No named filled-200k tok/s packet in this pass | Consumer GPU or Mac via the library build | Short session; raise num-ctx only after you watch VRAM |
| llama.cpp | Unsloth or bowmanslayer GGUF | 3090 fit cap 111,872 @ Q4_K_M; 5090 filled 262k 30.4 tok/s | 3090 24 GB; 5090 32 GB | Control, CUDA MTP, quantized KV experiments. Not one-card 200k on 24 GB |
| vLLM | Unsloth / InferAct / RadixArk NVFP4 (Blackwell) or BF16/FP8 | Spark recipes --max-model-len 262144 | DGX Spark 128 GB UMA | Serving with reserved native context on Spark/Blackwell |
Unsloth’s documented 27B paths (accessed 2026-09-10): llama.cpp + GGUF everywhere Ampere still lives; vllm serve unsloth/Qwen3.8-27B-NVFP4 on Blackwell, optional MTP speculative config. NVFP4 ~1.5× vs BF16 in Unsloth’s table is 1× B200, 128 concurrency — not a 3090 and not a 200k filled-decode number.
SGLang appears in Spark 27B recipes (amasu production is SGLang + DFlash2). It is not a fourth homepage runtime; it is the Spark serving fork of the same NVFP4 checkpoint.
Ollama: advertised 256K is not a 24 GB allocation
ollama run qwen3.8:27b is the ergonomic path. The library lists 18 GB and 256K. Unsloth’s 4-bit weight band is 16–19 GB. Those facts coexist with the KV table on the GGUF card: 200k f16 KV is ~12.5 GiB on top of weights. A 24 GB card does not get a free 256k because Ollama printed 256K.
Shortfall: no URL + date + Ollama version + 27B Q4 + hardware + filled 200k tok/s in this pass. Do not copy llama.cpp 41 tok/s onto Ollama. LocalRig’s only first-party Ollama figure is 8B on M4, 2026-06-27.
Ollama is also the library path for agent harnesses (ollama launch claude --model qwen3.8 and siblings on the same card). That does not allocate 200k. A coding agent that dumps a repo into context will OOM the same 24 GB card whether the wrapper is Ollama or llama.cpp. Fix the window, then the wrapper.
Unsloth Desktop is a convenience loader for the same GGUFs, not a fourth runtime in this comparison. If the binary cannot parse qwen35, update it; do not debug VRAM first. The library page is the Ollama packet; Hugging Face GGUF cards are the llama.cpp packet. Do not mix their tok/s.
Practical routing: use Ollama for a short session on the used 3090 cell. If the agent needs 200k, change silicon, then pick a runtime that has a reserved-262k recipe.
llama.cpp: the fit cap and the filled-context warning
llama.cpp is where the reproducible 24 GB cap lives. bowmanslayer/Qwen3.8-27B-GGUF: one RTX 3090, Vulkan, Q4_K_M, llama-fit-params max 111,872. CUDA no-draft 41.2 tok/s; CUDA MTP n-max 1 56.1 tok/s; Vulkan MTP is a rounding error. Those decode rows are 512-token greedy, not a 200k fill.
The filled-window warning is toki_mwc, 2026-08-19: RTX 5090 32 GB, Windows 11, llama.cpp b10430, UD-Q4_K_XL, 262,144 allocated, 28,167 MiB. Empty-context generation 63.7 tok/s. Filled 262k: prefill 4 m 27 s, generation 30.4 tok/s. Official FP8 weights left ~3.5 GiB and could not allocate 262k.
That is why this comparison is not “which runtime is faster.” Empty max-model-len and a repo-sized prompt are different jobs. Community-cited. Not LocalRig-measured.
CUDA build walkthrough (any GGUF, not 27B-specific): run llama.cpp on RTX 3090. Need qwen35 support; old binaries fail loudly.
vLLM: reserved 262k on Spark, not a 3090 GGUF clone
vLLM’s 27B long-context packets in September 2026 are NVFP4 or FP8 on Spark, with --max-model-len 262144 and usually --kv-cache-dtype fp8.
amasu/dgx-spark-qwen38: vLLM + NVFP4 + embedded MTP is the rollback stack; production moved to SGLang + DFlash2. Sample bench-matrix in that README: vLLM+MTP ~12–25 tok/s, DFlash2 ~32–56 tok/s on code/math. Context 262144. Agentic contexts noted around 110k with auto-truncate.
NVIDIA forum study (~2026-09-09): vLLM NVFP4+MTP TP=1 18.5 tok/s at C1; BF16 TP=1 4.5 tok/s. Bench is 128 in / 128 out with 262k reserved. Label it that way.
vLLM is still the right engine when several clients hit one 27B. That is the category-page constraint, and it still holds. At 200k the scheduler tax is real: every concurrent sequence pays KV. Spark recipes that reserve 262k with gpu_memory_utilization 0.45–0.80 are trading concurrency for window. Do not copy a C16 tok/s row onto a single long-running agent.
Multi-GPU wiring: vLLM multi-GPU. Hardware for the 200k reservation: Spark / Mac 128 / dual-3090. Dual-3090 split has no named filled-200k packet in this pass — capacity arithmetic only.
Unsloth: vllm serve unsloth/Qwen3.8-27B-NVFP4 plus optional --speculative-config MTP. Ampere is the wrong GPU for that command.
Decision table
| If you… | Runtime | Hardware |
|---|---|---|
| Want one command and a short session | Ollama qwen3.8:27b | Used 3090 / 4090 |
Want CUDA MTP, quantized KV, exact -c | llama.cpp | Same 24 GB card, cap ~112k @ Q4_K_M |
| Need a filled ~200k–262k agent on one box | vLLM or SGLang NVFP4 | Spark 128 GB (documented) or 32 GB+ llama.cpp GGUF (5090 packet) |
| Already own two 3090s and refuse Spark | llama.cpp or vLLM GGUF/HF split | 48 GB split; no filled-200k tok/s packet in this pass |
| Serve a team on one GPU | vLLM | Size for batch, not single-stream 41 tok/s |
Rent a Spark-class or 32 GB hour before buying if the only question is whether 200k prefill is livable: RunPod · Vast.ai · GPUMart.
Who this is NOT for
- Readers who want one tok/s winner. The packets are different cards, quants, and whether context was filled.
- 3090 owners expecting NVFP4. Wrong generation. GGUF only.
- People using the 2026-06-27 M4 8B row as a 27B 200k proxy. Different model.
- Flash-Next serving. Mia’s Spark Flash-Next runbooks are a different SKU. Flash selector.
- Anyone who needs LocalRig-measured 27B runtime numbers. Compile a public packet or run on owned silicon.
Methodology
- Ollama: library card accessed 2026-09-10. No filled-200k tok/s.
- llama.cpp 24 GB: bowmanslayer GGUF card, Vulkan fit + CUDA/Vulkan MTP tables.
- llama.cpp 32 GB 262k: toki_mwc 2026-08-19, b10430, UD-Q4_K_XL.
- vLLM / SGLang Spark: amasu README (accessed 2026-09-10) and NVIDIA forum 128/128 study with max-model-len 262144 (~2026-09-09).
- Unsloth NVFP4 vs BF16: vendor B200 table, not a 200k agent.
- LocalRig first-party: Llama 3.1 8B Q4_K_M, base M4, 2026-06-27 only. Not used as a 27B figure.
Sources
Frequently Asked Questions
Which runtime should I use for Qwen 3.8 27B at 200k?
llama.cpp GGUF on Ampere if you want a measured fit cap (one 3090 maxes around 112k). vLLM or SGLang NVFP4 on Spark/Blackwell if you are reserving 262k. Ollama ships an 18 GB 27b with a 256K advertised window; that is not a 24 GB 200k guarantee. LocalRig has no 27B first-party runtime bake-off.
Can Ollama allocate 256k on a 3090?
The library advertises 256K. A 24 GB card does not have a public llama-fit-params-style 200k Ollama packet in this pass. Treat advertised native context as the model card, not as your allocated num-ctx.
Is vLLM faster than llama.cpp for one user at 200k?
Not as a single number. Spark vLLM rows in the September 2026 forum study are 128/128 decode with 262k reserved. toki_mwc's llama.cpp 5090 row is filled 262k at 30.4 tok/s. Different hardware, quant, and whether the context was empty. Do not average them.
Does LocalRig's M4 Ollama vs llama.cpp test apply here?
No. That 2026-06-27 row is Llama 3.1 8B Q4_K_M. It is useful for the category page. It is not a 27B 200k packet.
Can I run NVFP4 on a 3090?
Unsloth's NVFP4 path requires Blackwell (RTX 50-series, DGX Spark, B200). Ampere 3090 stays on GGUF in llama.cpp or Ollama.