What Can I Run?

Can I Run Qwen 3.8 Locally? 27B vs Flash-Next vs the Giant MoE

Used RTX 3090 24GB for Qwen3.8-27B Q4
Top Pick Used RTX 3090 24GB for Qwen3.8-27B Q4

Short answer: run Qwen3.8-27B at Unsloth Dynamic 4-bit on a 24 GB card (used RTX 3090 is the community buy). Treat Flash-Next as a different product: DGX Spark / Blackwell NVFP4 / large unified memory, not a 3090 default. Do not buy a 2.4T MoE for a home box.

Qwen 3.8 is the current Qwen open family as of August 2026 (Qwen blog). The homepage instrument on LocalRig already uses 27B as the 24 GB floor because that is the dense SKU that actually maps to a used-card purchase. The last two weeks of X traffic mixed three names in the same sentence — 27B, Flash, Flash-Next — which is how people buy the wrong silicon.

This page is a constraint selector, not an Intelligence Index. LocalRig has no first-party Qwen 3.8 benchmark row. Every speed number is labeled with its source and date.

Which Qwen 3.8 SKU am I actually downloading?

Answer first: pick the SKU, then the card. Mixing names is the failure mode.

SKUWhat it isMemory class (vendor / community)LocalRig buy path
Qwen3.8-27BDense 27B vision-language model; 262k native contextUnsloth 4-bit: 16–19 GB RAM/VRAM (docs, 2026-08-19); UD-Q4_K_XL file ~17.6 GB on HF24 GB used 3090 / 4090. Headroom for KV cache and MTP.
Qwen3.8-Flash125B-class multimodal MoE (Qwen post, 2026-08-26)Unsloth Flash-Next GGUF: ~75 GB RAM at 1-bit; 96 GB machine preferred (docs)Not a 24 GB card. See Can I run Qwen 3.8 Flash locally?.
Qwen3.8-Flash-NextSparse / NVFP4 serving SKUSpark: Mia vLLM runbooks (~47 tok/s prose, 2026-09-06). Unsloth GGUF: ~78 GB unified/RAM (guide)DGX Spark or large unified memory. Blackwell NVFP4 on 50-series / Spark, not Ampere 3090.
Qwen3.8-2.4T-A95B2.4T total / 95B activeUnsloth 1-bit XXXS 397 GB; BF16 4.9 TBDo not buy home hardware for this. Rent or skip.

A September 2026 community map that matches this table, not a LocalRig ranking: 3090 / 4090 / 5090 / RTX PRO 6000 → 27B; one DGX Spark → Flash-Next (albustime, 2026-09-06).

Can a 24 GB card run Qwen 3.8 27B?

Yes, at 4-bit, with context discipline. Unsloth’s hardware table (accessed 2026-09-06) puts 4-bit Qwen3.8-27B at 16–19 GB of total memory, 6-bit at 23–26 GB, and BF16 at 56 GB. MTP (multi-token prediction) wants another 1–2 GB. That is why LocalRig still calls 24 GB the buy floor, not 16 GB: a 16 GB 5060-class card can load a 4-bit file and then lose the context window.

Weights on disk (Hugging Face file sizes, Unsloth Dynamic GGUF, accessed 2026-09-06):

Quant (Unsloth UD)Approx. fileFits 12 GB?Fits 24 GB with KV headroom?
UD-IQ1_S6.19 GBTight / emergencyYes, quality is the constraint
UD-Q3_K_XL13.1 GBBarely, almost no cacheYes
UD-Q4_K_XL17.6 GBNoYes — default
UD-Q5_K_M19.8 GBNoYes, less cache
Q8_029 GBNoNo on a single 24 GB card
BF1654.7 GBNoNo

Use the VRAM calculator for your context length. Unsloth’s 4-bit range is a vendor estimate, not a LocalRig lab measurement.

Browse used RTX 3090 24GB on eBay → · RTX 3090 listings on Amazon → · RTX 4090 on Amazon → · Used 3090 buying guide →

If you do not want to own a card yet, rent a 24 GB-class GPU and load the same GGUF: RunPod · Vast.ai · GPUMart 4090 monthly. Those links are affiliate-tagged; they do not change which SKU fits.

What public 27B decode numbers exist (not LocalRig)?

Named community and vendor-adjacent runs, each with hardware + runtime. None of these is a LocalRig measurement. Do not average them. Do not put them on the homepage as “our tok/s.”

SourceDateHardwareRuntime / notesDecode (as published)
sudoingX2026-08-30RTX 3090llama.cpp, MTP flag41 tok/s
lingster/qwen38-mtpaccessed 2026-09-07RTX 3090Q4_K_M, --spec-type draft-mtp31.0 → 41.3 tok/s
sudoingX paired table2026-08-17RTX 4090llama.cpp MTP on/off36 → 75 tok/s
same table2026-08-173×3090 tensor parallelllama.cpp MTP on/off49 → 96 tok/s
same table2026-08-17RTX 5090llama.cpp MTP on/off66 → 144 tok/s
analogalok2026-08-14RTX 409027B, MTP65 tok/s decode
analogalok Dflash22026-08-19RTX 4090custom llama.cpp patch, not stock Unsloth90 tok/s

Compilations that aggregate some of the same posts (third-party, accessed 2026-09-07 — still not LocalRig): vramcalculator.com Qwen3.8-27B speed (lists analogalok 4090 Q4_K_XL MTP 65 / no MTP 40.7, and the lingster 3090 pair) and Context Studios hardware guide (compiled 3090/4090/5090 prefill/decode table). Treat those as indexes into named configs.

A Medium Data Science Collective post argues 27B can look faster on a 3090 than a 4090 when the 3090 path is a tuned vLLM stack (~114 tok/s in that writeup) and the 4090 path is stock llama.cpp (~49 tok/s). That is a runtime comparison, not proof that Ampere beats Ada. Label the stack.

Unsloth’s own 27B NVFP4 batch table is on B200, not a 3090. Ampere buyers stay on GGUF.

Is Flash-Next worth a different machine?

Only if you already want Spark, 96 GB RAM, or 128 GB unified memory. Flash is not the SKU that justifies a used 3090. Full Flash / 75 GB GGUF selector: Can I run Qwen 3.8 Flash locally?.

Mia’s single-Spark vLLM NVFP4 numbers (community, not LocalRig):

  • 2026-09-04: NVFP4 on one Spark, 1M context, 37 tok/s prose single-stream (post).
  • 2026-09-05: 46 tok/s prose single-stream, 108 tok/s across 4 streams (post).
  • 2026-09-06: ~47 tok/s prose; 8k prefill +24.7% (post). Dual-Spark update the same week: 54 tok/s prose single-stream (post).

Repos to cite, not to copy as first-party: single Spark, dual Spark.

Unsloth’s NVFP4 27B path is a different stack again: Blackwell GPUs (RTX 50-series, DGX Spark, B200). Unsloth’s own batch table is on B200, not a 3090 — ~1.5× vs BF16 in that document. Ampere 3090 buyers should stay on GGUF, not NVFP4.

NVIDIA DGX Spark on Amazon → · Spark 2-pack → · Full ownership comparison: DGX Spark vs 2× RTX 3090.

How do I actually load 27B without guessing the runtime?

Start with Unsloth’s documented GGUF + a current llama.cpp, or Unsloth Desktop. Qwen3.8-27B uses a Qwen3.5-family architecture (qwen35 on the GGUF card). Old Ollama builds will fail loudly; that is a runtime problem, not a VRAM problem.

Unsloth’s documented 27B command (vendor docs, accessed 2026-09-06):

hf download unsloth/Qwen3.8-27B-GGUF \
    --local-dir unsloth/Qwen3.8-27B-GGUF \
    --include "*UD-Q4_K_XL*"

./llama.cpp/llama-cli \
    --model unsloth/Qwen3.8-27B-GGUF/Qwen3.8-27B-UD-Q4_K_XL.gguf \
    --temp 1.0 --top-p 0.95 --top-k 20 --min-p 0.0

Thinking vs instruct sampling differs (Unsloth table): thinking temp=1.0 / top_p=0.95; instruct temp=0.7 / top_p=0.80 plus presence_penalty=1.5. Native context is 262,144 tokens; 1M is YaRN, not free headroom on 24 GB.

Hermes Desktop one-click for Unsloth GGUFs including Qwen3.8-27B was announced 2026-09-03 (Unsloth). That is a convenience path, not a reason to skip checking file size against VRAM.

Intersection cells (cite-packet pages, not this SKU overview): 27B Q4 on a used 3090 at 4k–32k · 27B Q4 at 200k: Spark / Mac 128 / dual 3090 · Ollama vs llama.cpp vs vLLM at 200k.

Related runtime guides: Ollama vs llama.cpp vs vLLM · run llama.cpp on RTX 3090.

Should I believe “Flash-Next beats GPT-5.6” posts?

Not as a buying input. Those posts are answering an Artificial Analysis / frontier-index question. LocalRig’s question is whether the weights fit and whether the runtime you can install will load them. If you want a quality ranking, read the index publisher and the named X threads; do not treat this page as a substitute. A high-reach example from this week: analogalok on AA v4.2, 2026-09-05. Ahmad Osman’s personal stack the week of 2026-08-27 put GLM-5.3-Flash above Flash-Next above 27B (post) — that is preference, not a VRAM table.

What about GLM-5.3-Flash in the same feed?

It is a different family. Unsloth’s GLM-5.3-Flash GGUF path (320B-A18B / ox-alpha) is documented as 1-bit ~100 GB and 3-bit on 128 GB setups. That is Mac Studio / Spark / multi-GPU rent, not a 3090. Selector: Can I run GLM-5.3-Flash locally?. LocalRig’s older GLM-5.2 page is about a larger flagship that still does not fit home cards. Do not collapse 5.2 and 5.3-Flash.

The 24 GB GLM Flash in the same batch is GLM-4.7-Flash (~18 GB 4-bit), not 5.3. Other 24 GB guests: Gemma 4 31B, gpt-oss-20b, Ornith 9B, Qwen 3.6-27B. Skip/rent (do not size a 3090): Kimi K2.6, Kimi K3, DeepSeek-V4-Pro. Full lanes: Can I run it? index.

Who this is NOT for

  • Anyone who wants LocalRig tok/s on Qwen 3.8. We do not have a lab row yet. If you need a measured number, run it yourself or cite a named config.
  • Buyers shopping Flash-Next on a 12–24 GB Ampere card. Wrong SKU. Buy 27B GGUF or change the machine.
  • People trying to run Qwen3.8-2.4T at home. Unsloth’s 1-bit file is hundreds of gigabytes. Rent or skip.
  • Qwen3.5-only runtimes. Use the Qwen3.5 size picker until you can load 3.8.
  • Production multi-user serving on a single 24 GB card. Fit is not a scheduler. See vLLM multi-GPU.

Methodology

  • Fit numbers: Hugging Face GGUF file sizes and Unsloth’s published RAM/VRAM bands, accessed 2026-09-07. Labeled as vendor estimates.
  • 27B decode numbers: Named X posts and the lingster GitHub recipe listed in the public-speed table. Community / operator-reported. Not LocalRig-measured. Hardware, quant, and runtime stay on the row. Compilations (vramcalculator, Context Studios) are indexes, not first-party data.
  • Spark decode numbers: Mia X posts 2026-09-04–2026-09-06 plus GitHub runbooks. Community / operator-reported, not LocalRig-measured. Config is Spark + vLLM NVFP4 unless a post says otherwise.
  • SKU identities: Qwen blog (August 2026), Unsloth model cards, Qwen’s 2026-08-26 Flash announcement quoted by Unsloth.
  • LocalRig first-party: none for Qwen 3.8 as of 2026-09-07. No lab hour; no engine.db row.

Sources

Frequently Asked Questions

Can I run Qwen 3.8 27B on a used RTX 3090?

Yes at a 4-bit Unsloth Dynamic GGUF. Unsloth's 2026-08-19 table lists 4-bit Qwen3.8-27B at 16–19 GB of RAM/VRAM before long context and MTP headroom. A 24 GB 3090 is the LocalRig buy floor because it leaves KV-cache room. LocalRig has not measured tok/s on this card.

Is Qwen 3.8 Flash-Next the same as Qwen 3.8 27B?

No. 27B is a dense vision-language model. Flash / Flash-Next GGUF is a 125B MoE sized around 75 GB at 1-bit. Spark NVFP4 is a third stack. See the Flash selector for that SKU; do not buy a 24 GB card expecting Flash to behave like 27B.

What hardware should I buy if I want Flash-Next specifically?

A DGX Spark (or two) if you want Mia's published vLLM NVFP4 runbooks, or a 128 GB-class Mac / RAM box if you are following Unsloth's GGUF path. Rent a Blackwell or Spark-class GPU first if you have not confirmed the stack.

Should I still download Qwen3.5 instead?

Only if your runtime cannot load Qwen3.8 yet. The current dense daily driver on 24 GB cards in September 2026 community posts is Qwen3.8-27B, not Qwen3.5-32B. Keep the Qwen3.5 guide for older stacks.

Does LocalRig have first-party Qwen 3.8 tok/s?

No. Speed figures below are Unsloth vendor tables or named X/GitHub runbooks, dated, with hardware and runtime labeled. They are not independently verified by LocalRig and are not a lab row in engine.db.

Sources

  • Qwen Team, Qwen3.8-Max blog, https://qwen.ai/blog?id=qwen3.8, August 2026, accessed 2026-09-06
  • unsloth/Qwen3.8-27B-GGUF model card, https://huggingface.co/unsloth/Qwen3.8-27B-GGUF, accessed 2026-09-06
  • Unsloth, How to Run Qwen3.8 Locally, https://unsloth.ai/docs/models/qwen3.8, Dynamic V3.0 note 2026-08-19, accessed 2026-09-07
  • Unsloth, Qwen3.8-Flash-Next local guide, https://unsloth.ai/docs/models/qwen3.8-next, accessed 2026-09-07
  • Mia (@MiaAI_lab) DGX Spark Flash-Next numbers, 2026-09-04 through 2026-09-06, X posts cited in body
  • MiaAI-Lab Spark vLLM runbooks, https://github.com/MiaAI-Lab/Qwen3.8-Flash-Next-Single-DGX-Spark, accessed 2026-09-06
  • sudoingX, 27B on RTX 3090 llama.cpp MTP 41 tok/s, 2026-08-30, https://x.com/sudoingX/status/2094081818853298653
  • sudoingX paired MTP table, 2026-08-17, https://x.com/sudoingX/status/2089227419765080440
  • lingster/qwen38-mtp, 3090 Q4_K_M 31.0 → 41.3 tok/s, accessed 2026-09-07
  • analogalok, 4090 27B MTP 65 tok/s, 2026-08-14, https://x.com/analogalok/status/2088326480669667699