What Can I Run?

Can I Run Qwen3.6-27B Locally? Prior 24 GB Driver, Not Qwen 3.8

Used RTX 3090 24GB for Unsloth 4-bit (18 GB); 6-bit is the 24 GB ceiling
Top Pick Used RTX 3090 24GB for Unsloth 4-bit (18 GB); 6-bit is the 24 GB ceiling

Short answer: Qwen/Qwen3.6-27B is Alibaba’s prior 27B multimodal hybrid-thinking SKU (Hub: ~27.78B, Apache-2.0, architecture qwen3_5). Unsloth (accessed 2026-09-10): 4-bit 18 GB, 6-bit 24 GB, 8-bit 30 GB, BF16 55 GB. A used RTX 3090 is the 4-bit cell. Qwen 3.8 27B is the current LocalRig homepage dense default — different weights. Unsloth MTP ~160 tok/s is on an RTX 6000, not a 3090. LocalRig has not measured 3.6 on Ampere.

3.6 vs 3.8 vs Flash-Next

Three different downloads. Qwen 3.8 27B is the September 2026 dense daily driver on 24 GB Q4, with hybrid-attention KV notes on the 200k page. Flash-Next is the 75 GB GGUF class. 3.6-27B is the older 27B Unsloth still documents at 18 GB 4-bit. Hub ~26M downloads (2026-09-10) is volume, not a reason to ignore 3.8.

35B-A3B is a different 3.6 row (Unsloth 4-bit 23 GB). This page is 27B only.

Unsloth memory table

Copied 2026-09-10. Units = RAM + VRAM or unified. 27B row:

3-bit4-bit6-bit8-bitBF16
15 GB18 GB24 GB30 GB55 GB

Unsloth intro: “Qwen3.6-27B runs on 18GB RAM setups.” MTP adds ~1 GB (27B MTP 4-bit 19 GB). NVFP4 (Jul 10, 2026) is a Blackwell / Spark path; Unsloth says GGUF for older GPUs. Do not mix NVFP4 Spark recipes with Ampere GGUF.

Context: 256K in the llama.cpp notes. Filled 256k on a 3090 is a shortfall unless a cite packet appears.

Used 3090 vs 16 GB

24 GB matches Unsloth 4-bit 18 GB with KV room; 6-bit 24 GB is the ceiling. A 16 GB card is under the 18 GB 4-bit band. 3-bit 15 GB is the 16 GB maybe — quality trade, not LocalRig’s Q4 cell.

Browse used RTX 3090 24GB on eBay →

Check RTX 4060 Ti 16GB on Amazon → only for the 3-bit band, not 4-bit 18 GB plus a long window.

Rent: RunPod, Vast.ai, GPUMart.

If you are choosing one 27B for a new 3090 build in September 2026, LocalRig’s current selector is 3.8, not 3.6. Keep 3.6 if you already downloaded it and it fits Unsloth’s 18 GB 4-bit band.

MTP speed — what you may quote

Unsloth: Qwen3.6 27B MTP ~160 tok/s on an RTX 6000; 35B-A3B 240 tok/s on the same class. Community-cited, RTX 6000 only. --spec-draft-n-max 2 is their usual starting point. Not a 3090 cell. llama.cpp MTP is documented as merged; still not a LocalRig first-party row.

Related: Qwen 3.8 27B Q4 on used 3090. 200k hardware is 3.8, not 3.6.

Why LocalRig still points new 3090 buyers at 3.8, not 3.6

3.6 is documented; 3.8 is the current dense default. Unsloth’s 3.6 4-bit band (18 GB) and 6-bit (24 GB) make a used 3090 a valid 3.6 Q4 machine. That does not make 3.6 the homepage SKU. Qwen 3.8 27B has the 4k–32k 3090 cell and the 200k Spark/Mac/dual-3090 cell. Those pages are 3.8 hybrid-attention math. Do not copy 3.8 KV footnotes onto 3.6.

MTP: Unsloth ~160 tok/s on RTX 6000 for 27B MTP, ~1 GB extra VRAM. Ampere 3090 GGUF without MTP is a different speed class — unmeasured here. NVFP4 2.5× (Jul 10, 2026) is Blackwell/Spark. A used 3090 stays on GGUF.

256K context: Unsloth llama.cpp notes name the max. A 24 GB card at 6-bit (24 GB band) has no spare for a filled 256k KV. Q4 at 18 GB is the 3090 cell with room. 16 GB is Unsloth 3-bit (15 GB) or offload.

35B-A3B (Unsloth 4-bit 23 GB) is the other 3.6 SKU. It is tighter on one 3090 than 27B Q4. This URL is 27B only.

Hub ~26M downloads do not outrank 3.8. If you already have 3.6 GGUFs on disk and Unsloth’s 18 GB 4-bit fits, keep using them. If you are starting a 24 GB library in September 2026, start at 3.8.

Used RTX 3090 buying guide. Ollama vs llama.cpp vs vLLM.

Unsloth’s KL/PPL table for 3.6-27B (4-bit 26.2 GB in that quality table vs 18 GB in the hardware table) is a reminder that two Unsloth tables can disagree in presentation. This page uses the hardware requirements table (4-bit 18 GB, 6-bit 24 GB) as the buy constraint. The GGUF benchmark table’s 26.2 GB 4-bit line is a different measurement context. Do not average them. If your download’s file size is closer to 26 GB, you are in the 6-bit/24 GB conversation on a 3090, not the 18 GB 4-bit band.

MLX: Unsloth documents mlx_vlm.chat for 3.6-27B 4-bit. That is a Mac path, not a 3090 GGUF path, and not a LocalRig Metal tok/s. Best Mac for local LLM if unified memory is the purchase — 32 GB unified is Unsloth 6-bit 24 GB on the edge; 18 GB 4-bit is the more honest Mac SKU.

Do not follow 3.8 200k dual-3090 math for 3.6. Different weights, different attention notes.

Thinking vs non-thinking on 3.6 still costs KV when thinking is on. Unsloth Studio auto-sets sampling; llama.cpp users own --temp and chat-template flags. LocalRig is not pasting Unsloth Studio loopback URLs.

If your 3.6 GGUF is an uncensored third-party merge, that is a different repo. This page sizes Qwen/Qwen3.6-27B + Unsloth GGUF only.

Qwen 3.8 27B remains the LocalRig dense default on 24 GB because that is the current daily-driver SKU, not because 3.6 “failed.” If you already have a 3.6 GGUF on disk, Unsloth’s 18 GB 4-bit / 24 GB 6-bit table still governs the 3090. Do not mix 3.8 200k KV notes or Flash-Next 75 GB files into this download. MTP ~160 tok/s stays on RTX 6000 in Unsloth’s copy; a used 3090 does not inherit it.

Who this is NOT for

  • Readers who meant Qwen 3.8 27B or Flash-Next. Different slugs, different files.
  • 16 GB 4-bit shoppers. 18 GB band + KV.
  • Anyone quoting 160 tok/s on a 3090. That Unsloth figure is RTX 6000 MTP.
  • 35B-A3B shoppers. Other Unsloth row (23 GB 4-bit).
  • People who need a LocalRig 3.6 first-party bench. None as of 2026-09-10.

Methodology

  • Fit numbers: Unsloth Qwen3.6 table, 2026-09-10 (18 GB 4-bit, 24 GB 6-bit).
  • SKU identity: Qwen/Qwen3.6-27B Hub MCP (~27.78B).
  • Speed: Unsloth MTP 160 tok/s on RTX 6000 labeled vendor; not transplanted.
  • LocalRig first-party: none for 3.6-27B as of 2026-09-10.

Sources

Frequently Asked Questions

Is Qwen3.6-27B the same as Qwen 3.8 27B?

No. 3.8 is the current LocalRig dense daily driver. 3.6 is the prior 27B SKU. Do not reuse 3.8 hybrid-KV notes or Flash-Next files on 3.6.

Can a used 3090 run Qwen3.6-27B Q4?

Unsloth's 4-bit band is 18 GB total memory. A 24 GB 3090 matches that class. 6-bit is 24 GB — tight once KV starts. LocalRig has no first-party tok/s for 3.6 on 3090.

What about Unsloth's 160 tok/s figure?

That MTP number is Qwen3.6-27B on an RTX 6000 in Unsloth's docs. It is not a 3090, Spark, or Mac figure. Do not transplant it.

16 GB card?

Unsloth 3-bit is 15 GB; 4-bit is 18 GB. A 16 GB card is a 3-bit conversation or offload, not the 4-bit 18 GB band plus KV.

Flash-Next?

Qwen 3.8 Flash-Next is a different, larger GGUF. See that selector. 3.6-27B is the 24 GB 27B SKU from the 3.6 family.

Sources

  • Unsloth Qwen3.6 local guide, https://unsloth.ai/docs/models/qwen3.6, accessed 2026-09-10
  • Qwen/Qwen3.6-27B, https://huggingface.co/Qwen/Qwen3.6-27B, Hub MCP 2026-09-10
  • unsloth/Qwen3.6-27B-GGUF, https://huggingface.co/unsloth/Qwen3.6-27B-GGUF