Can I Run Qwen 3.8 Locally? 27B vs Flash-Next vs the Giant MoE
Short answer: run Qwen3.8-27B at Unsloth Dynamic 4-bit on a 24 GB card (used RTX 3090 is the community buy). Treat Flash-Next as a different product: DGX Spark / Blackwell NVFP4 / large unified memory, not a 3090 default. Do not buy a 2.4T MoE for a home box.
Qwen 3.8 is the current Qwen open family as of August 2026 (Qwen blog). The homepage instrument on LocalRig already uses 27B as the 24 GB floor because that is the dense SKU that actually maps to a used-card purchase. The last two weeks of X traffic mixed three names in the same sentence — 27B, Flash, Flash-Next — which is how people buy the wrong silicon.
This page is a constraint selector, not an Intelligence Index. LocalRig has no first-party Qwen 3.8 benchmark row. Every speed number is labeled with its source and date.
Which Qwen 3.8 SKU am I actually downloading?
Answer first: pick the SKU, then the card. Mixing names is the failure mode.
| SKU | What it is | Memory class (vendor / community) | LocalRig buy path |
|---|---|---|---|
| Qwen3.8-27B | Dense 27B vision-language model; 262k native context | Unsloth 4-bit: 16–19 GB RAM/VRAM (docs, 2026-08-19); UD-Q4_K_XL file ~17.6 GB on HF | 24 GB used 3090 / 4090. Headroom for KV cache and MTP. |
| Qwen3.8-Flash | 125B-class multimodal MoE (Qwen post, 2026-08-26) | Unsloth Flash-Next GGUF: ~75 GB RAM at 1-bit; 96 GB machine preferred (docs) | Not a 24 GB card. See Can I run Qwen 3.8 Flash locally?. |
| Qwen3.8-Flash-Next | Sparse / NVFP4 serving SKU | Spark: Mia vLLM runbooks (~47 tok/s prose, 2026-09-06). Unsloth GGUF: ~78 GB unified/RAM (guide) | DGX Spark or large unified memory. Blackwell NVFP4 on 50-series / Spark, not Ampere 3090. |
| Qwen3.8-2.4T-A95B | 2.4T total / 95B active | Unsloth 1-bit XXXS 397 GB; BF16 4.9 TB | Do not buy home hardware for this. Rent or skip. |
A September 2026 community map that matches this table, not a LocalRig ranking: 3090 / 4090 / 5090 / RTX PRO 6000 → 27B; one DGX Spark → Flash-Next (albustime, 2026-09-06).
Can a 24 GB card run Qwen 3.8 27B?
Yes, at 4-bit, with context discipline. Unsloth’s hardware table (accessed 2026-09-06) puts 4-bit Qwen3.8-27B at 16–19 GB of total memory, 6-bit at 23–26 GB, and BF16 at 56 GB. MTP (multi-token prediction) wants another 1–2 GB. That is why LocalRig still calls 24 GB the buy floor, not 16 GB: a 16 GB 5060-class card can load a 4-bit file and then lose the context window.
Weights on disk (Hugging Face file sizes, Unsloth Dynamic GGUF, accessed 2026-09-06):
| Quant (Unsloth UD) | Approx. file | Fits 12 GB? | Fits 24 GB with KV headroom? |
|---|---|---|---|
| UD-IQ1_S | 6.19 GB | Tight / emergency | Yes, quality is the constraint |
| UD-Q3_K_XL | 13.1 GB | Barely, almost no cache | Yes |
| UD-Q4_K_XL | 17.6 GB | No | Yes — default |
| UD-Q5_K_M | 19.8 GB | No | Yes, less cache |
| Q8_0 | 29 GB | No | No on a single 24 GB card |
| BF16 | 54.7 GB | No | No |
Use the VRAM calculator for your context length. Unsloth’s 4-bit range is a vendor estimate, not a LocalRig lab measurement.
Browse used RTX 3090 24GB on eBay → · RTX 3090 listings on Amazon → · RTX 4090 on Amazon → · Used 3090 buying guide →
If you do not want to own a card yet, rent a 24 GB-class GPU and load the same GGUF: RunPod · Vast.ai · GPUMart 4090 monthly. Those links are affiliate-tagged; they do not change which SKU fits.
What public 27B decode numbers exist (not LocalRig)?
Named community and vendor-adjacent runs, each with hardware + runtime. None of these is a LocalRig measurement. Do not average them. Do not put them on the homepage as “our tok/s.”
| Source | Date | Hardware | Runtime / notes | Decode (as published) |
|---|---|---|---|---|
| sudoingX | 2026-08-30 | RTX 3090 | llama.cpp, MTP flag | 41 tok/s |
| lingster/qwen38-mtp | accessed 2026-09-07 | RTX 3090 | Q4_K_M, --spec-type draft-mtp | 31.0 → 41.3 tok/s |
| sudoingX paired table | 2026-08-17 | RTX 4090 | llama.cpp MTP on/off | 36 → 75 tok/s |
| same table | 2026-08-17 | 3×3090 tensor parallel | llama.cpp MTP on/off | 49 → 96 tok/s |
| same table | 2026-08-17 | RTX 5090 | llama.cpp MTP on/off | 66 → 144 tok/s |
| analogalok | 2026-08-14 | RTX 4090 | 27B, MTP | 65 tok/s decode |
| analogalok Dflash2 | 2026-08-19 | RTX 4090 | custom llama.cpp patch, not stock Unsloth | 90 tok/s |
Compilations that aggregate some of the same posts (third-party, accessed 2026-09-07 — still not LocalRig): vramcalculator.com Qwen3.8-27B speed (lists analogalok 4090 Q4_K_XL MTP 65 / no MTP 40.7, and the lingster 3090 pair) and Context Studios hardware guide (compiled 3090/4090/5090 prefill/decode table). Treat those as indexes into named configs.
A Medium Data Science Collective post argues 27B can look faster on a 3090 than a 4090 when the 3090 path is a tuned vLLM stack (~114 tok/s in that writeup) and the 4090 path is stock llama.cpp (~49 tok/s). That is a runtime comparison, not proof that Ampere beats Ada. Label the stack.
Unsloth’s own 27B NVFP4 batch table is on B200, not a 3090. Ampere buyers stay on GGUF.
Is Flash-Next worth a different machine?
Only if you already want Spark, 96 GB RAM, or 128 GB unified memory. Flash is not the SKU that justifies a used 3090. Full Flash / 75 GB GGUF selector: Can I run Qwen 3.8 Flash locally?.
Mia’s single-Spark vLLM NVFP4 numbers (community, not LocalRig):
- 2026-09-04: NVFP4 on one Spark, 1M context, 37 tok/s prose single-stream (post).
- 2026-09-05: 46 tok/s prose single-stream, 108 tok/s across 4 streams (post).
- 2026-09-06: ~47 tok/s prose; 8k prefill +24.7% (post). Dual-Spark update the same week: 54 tok/s prose single-stream (post).
Repos to cite, not to copy as first-party: single Spark, dual Spark.
Unsloth’s NVFP4 27B path is a different stack again: Blackwell GPUs (RTX 50-series, DGX Spark, B200). Unsloth’s own batch table is on B200, not a 3090 — ~1.5× vs BF16 in that document. Ampere 3090 buyers should stay on GGUF, not NVFP4.
NVIDIA DGX Spark on Amazon → · Spark 2-pack → · Full ownership comparison: DGX Spark vs 2× RTX 3090.
How do I actually load 27B without guessing the runtime?
Start with Unsloth’s documented GGUF + a current llama.cpp, or Unsloth Desktop. Qwen3.8-27B uses a Qwen3.5-family architecture (qwen35 on the GGUF card). Old Ollama builds will fail loudly; that is a runtime problem, not a VRAM problem.
Unsloth’s documented 27B command (vendor docs, accessed 2026-09-06):
hf download unsloth/Qwen3.8-27B-GGUF \
--local-dir unsloth/Qwen3.8-27B-GGUF \
--include "*UD-Q4_K_XL*"
./llama.cpp/llama-cli \
--model unsloth/Qwen3.8-27B-GGUF/Qwen3.8-27B-UD-Q4_K_XL.gguf \
--temp 1.0 --top-p 0.95 --top-k 20 --min-p 0.0
Thinking vs instruct sampling differs (Unsloth table): thinking temp=1.0 / top_p=0.95; instruct temp=0.7 / top_p=0.80 plus presence_penalty=1.5. Native context is 262,144 tokens; 1M is YaRN, not free headroom on 24 GB.
Hermes Desktop one-click for Unsloth GGUFs including Qwen3.8-27B was announced 2026-09-03 (Unsloth). That is a convenience path, not a reason to skip checking file size against VRAM.
Intersection cells (cite-packet pages, not this SKU overview): 27B Q4 on a used 3090 at 4k–32k · 27B Q4 at 200k: Spark / Mac 128 / dual 3090 · Ollama vs llama.cpp vs vLLM at 200k.
Related runtime guides: Ollama vs llama.cpp vs vLLM · run llama.cpp on RTX 3090.
Should I believe “Flash-Next beats GPT-5.6” posts?
Not as a buying input. Those posts are answering an Artificial Analysis / frontier-index question. LocalRig’s question is whether the weights fit and whether the runtime you can install will load them. If you want a quality ranking, read the index publisher and the named X threads; do not treat this page as a substitute. A high-reach example from this week: analogalok on AA v4.2, 2026-09-05. Ahmad Osman’s personal stack the week of 2026-08-27 put GLM-5.3-Flash above Flash-Next above 27B (post) — that is preference, not a VRAM table.
What about GLM-5.3-Flash in the same feed?
It is a different family. Unsloth’s GLM-5.3-Flash GGUF path (320B-A18B / ox-alpha) is documented as 1-bit ~100 GB and 3-bit on 128 GB setups. That is Mac Studio / Spark / multi-GPU rent, not a 3090. Selector: Can I run GLM-5.3-Flash locally?. LocalRig’s older GLM-5.2 page is about a larger flagship that still does not fit home cards. Do not collapse 5.2 and 5.3-Flash.
The 24 GB GLM Flash in the same batch is GLM-4.7-Flash (~18 GB 4-bit), not 5.3. Other 24 GB guests: Gemma 4 31B, gpt-oss-20b, Ornith 9B, Qwen 3.6-27B. Skip/rent (do not size a 3090): Kimi K2.6, Kimi K3, DeepSeek-V4-Pro. Full lanes: Can I run it? index.
Who this is NOT for
- Anyone who wants LocalRig tok/s on Qwen 3.8. We do not have a lab row yet. If you need a measured number, run it yourself or cite a named config.
- Buyers shopping Flash-Next on a 12–24 GB Ampere card. Wrong SKU. Buy 27B GGUF or change the machine.
- People trying to run Qwen3.8-2.4T at home. Unsloth’s 1-bit file is hundreds of gigabytes. Rent or skip.
- Qwen3.5-only runtimes. Use the Qwen3.5 size picker until you can load 3.8.
- Production multi-user serving on a single 24 GB card. Fit is not a scheduler. See vLLM multi-GPU.
Methodology
- Fit numbers: Hugging Face GGUF file sizes and Unsloth’s published RAM/VRAM bands, accessed 2026-09-07. Labeled as vendor estimates.
- 27B decode numbers: Named X posts and the lingster GitHub recipe listed in the public-speed table. Community / operator-reported. Not LocalRig-measured. Hardware, quant, and runtime stay on the row. Compilations (vramcalculator, Context Studios) are indexes, not first-party data.
- Spark decode numbers: Mia X posts 2026-09-04–2026-09-06 plus GitHub runbooks. Community / operator-reported, not LocalRig-measured. Config is Spark + vLLM NVFP4 unless a post says otherwise.
- SKU identities: Qwen blog (August 2026), Unsloth model cards, Qwen’s 2026-08-26 Flash announcement quoted by Unsloth.
- LocalRig first-party: none for Qwen 3.8 as of 2026-09-07. No lab hour; no
engine.dbrow.
Sources
- Qwen3.8 blog — August 2026 family announcement.
- unsloth/Qwen3.8-27B-GGUF — file sizes and architecture tag.
- Unsloth Qwen3.8 run guide — 4-bit 16–19 GB table; Dynamic V3.0 2026-08-19.
- Unsloth Qwen3.8-Flash-Next guide and Flash-Next GGUF.
- unsloth/Qwen3.8-27B-NVFP4 — Blackwell-only path.
- Mia Spark posts and single-Spark runbook / dual-Spark runbook.
- X daily-driver map: albustime, 2026-09-06.
- Public 27B speed table: sudoingX 2026-08-30 / 2026-08-17, lingster/qwen38-mtp, analogalok 2026-08-14 (MTP) and 2026-08-19 (Dflash2 custom patch).
- Flash 75 GB selector: Can I run Qwen 3.8 Flash locally?.
- September SKU index: Can I run it?.
- Ingest log:
knowledge/research/x-ingest-2026-09-06.md.
Frequently Asked Questions
Can I run Qwen 3.8 27B on a used RTX 3090?
Yes at a 4-bit Unsloth Dynamic GGUF. Unsloth's 2026-08-19 table lists 4-bit Qwen3.8-27B at 16–19 GB of RAM/VRAM before long context and MTP headroom. A 24 GB 3090 is the LocalRig buy floor because it leaves KV-cache room. LocalRig has not measured tok/s on this card.
Is Qwen 3.8 Flash-Next the same as Qwen 3.8 27B?
No. 27B is a dense vision-language model. Flash / Flash-Next GGUF is a 125B MoE sized around 75 GB at 1-bit. Spark NVFP4 is a third stack. See the Flash selector for that SKU; do not buy a 24 GB card expecting Flash to behave like 27B.
What hardware should I buy if I want Flash-Next specifically?
A DGX Spark (or two) if you want Mia's published vLLM NVFP4 runbooks, or a 128 GB-class Mac / RAM box if you are following Unsloth's GGUF path. Rent a Blackwell or Spark-class GPU first if you have not confirmed the stack.
Should I still download Qwen3.5 instead?
Only if your runtime cannot load Qwen3.8 yet. The current dense daily driver on 24 GB cards in September 2026 community posts is Qwen3.8-27B, not Qwen3.5-32B. Keep the Qwen3.5 guide for older stacks.
Does LocalRig have first-party Qwen 3.8 tok/s?
No. Speed figures below are Unsloth vendor tables or named X/GitHub runbooks, dated, with hardware and runtime labeled. They are not independently verified by LocalRig and are not a lab row in engine.db.