Homelab & Platform

How Do You Run Qwen 3.8 Flash-Next on Two DGX Sparks in 2026?

DGX Spark 2-pack with cable
Top Pick DGX Spark 2-pack with cable

Short answer: two NVIDIA DGX Spark units can serve Qwen3.8-Flash-Next as a two-node vLLM NVFP4 job if you follow Mia’s public dual-Spark runbook (accessed 2026-09-07): ConnectX between boxes, Docker on both nodes, and the NVFP4 checkpoint — not Unsloth’s 75 GB GGUF. Public prose decode from Mia is 54 tok/s single-stream on the dual kit (X, 2026-09-05) versus ~47 tok/s on one Spark (X, 2026-09-06). Two Sparks are not one 256 GB pool. LocalRig has no first-party tok/s for this stack.

This is a how-to for operators who already own, or are buying, a matching pair. Hardware-fit comparison against used 3090s lives at DGX Spark vs dual RTX 3090. The 27B dense SKU on 24 GB is a different article: Can I run Qwen 3.8 locally?.

Buy the DGX Spark 2-pack with cable on Amazon → · One DGX Spark →

What hardware and software does this stack actually need?

Two DGX Spark nodes, a ConnectX link, Docker, and Mia’s vLLM image — not a pair of 24 GB cards and not llama.cpp GGUF. The dual README (accessed 2026-09-07) lists prerequisites as two GB10 Sparks with 128 GB unified memory each (sm_121), connected via ConnectX RoCE/IB, passwordless SSH, Docker on both nodes, and about 126 GiB free on each node for the checkpoint. Default layout: each node keeps its own copy under ~/.cache/huggingface; start.sh rsyncs the worker copy from the head once.

The public checkpoint is RadixArk/Qwen3.8-Flash-Next-NVFP4. Mia’s README describes that NVIDIA NVFP4 checkpoint as 124 GiB on disk / 11 shards (NVFP4 routed experts, an FP8 PLE n-gram table on the order of 51 GB, bf16 attention/dense/vision). The serving image named in the same README is vllm/vllm-openai:qwen38-flash-next. Parallelism in the shipped .env.sample is tensor parallel size 2, expert parallel on, MTP speculative tokens = 3.

That is a cluster serving recipe. It is not “install Ollama and pull a tag.” If you only have one Spark, use the sibling runbook instead: Qwen3.8-Flash-Next-Single-DGX-Spark.

Is a dual Spark setup the same as 256GB of unified memory?

No. Two 128 GB appliances linked for vLLM TP2 are still two computers. The DGX Spark 2-pack with cable is a hardware and interconnect path. Whether a model uses memory across both units depends on the runtime and the parallelism mode. It is not automatically equivalent to one 256 GB unified-memory machine.

Mia’s dual README is explicit about how TP2 spends the boxes. Weights land per GPU: the KV-cache section reports weights at 64.46 GiB on the running worker (68.52 GiB including non-torch) at GPU_MEMORY_UTILIZATION=0.835, with a 32.02 GiB KV pool. The same section states that TP2 shards each token across both GPUs, so you should not add the two 128 GB numbers and call the result a single pool. The useful token space Mia published for the default fp8 KV config is 3,652,200 cache tokens on that kit — one shared block pool, not 3.65M plus another 3.65M.

Buy the 2-pack when you already want this two-node serving plan and matching systems plus the cable. Start from one Spark if you have not confirmed that the interconnect, Docker image, and checkpoint load on a single node first.

Is this the same as Unsloth’s 75GB Flash GGUF?

No. Dual-Spark NVFP4 and Unsloth Flash-Next GGUF are different stacks that share a family name. Unsloth’s 1-bit GGUF is about 75 GB of total memory on a 96 GB RAM or unified-memory machine; that selector is Can I run Qwen 3.8 Flash locally?. Mia’s published Spark numbers are vLLM NVFP4 on GB10. Do not paste Unsloth’s RTX PRO 6000 MTP column into a Spark table, and do not paste Mia’s 54 tok/s into a llama.cpp GGUF table.

Name you sawRuntime / quantMachine classLocalRig page
Qwen3.8-27B Q4llama.cpp / GGUF on 24 GBUsed 3090 / 409027B selector — not this SKU
Flash GGUF / Unsloth Flash-Nextllama.cpp or Unsloth Desktop, 75 GB at 1-bit96 GB RAM / unifiedFlash GGUF selector
Flash-Next on 1× SparkvLLM NVFP4One DGX SparkMia single-Spark repo
Flash-Next on 2× SparkvLLM NVFP4, TP2+EPTwo Sparks + ConnectXThis page

If your actual download is the 27B dense file, stop. Dual 3090s (48 GB split) are the cheaper path for that SKU and the wrong path for this NVFP4 checkpoint.

What public throughput numbers exist, and whose are they?

Mia’s, dated, on named kits — not LocalRig lab hours. Treat every tok/s below as a cited operator config. Hardware, runtime, and date travel with the number.

ClaimHardwareRuntime / configSourceDate
54 tok/s prose, single-streamDual DGX SparkMia dual update (vLLM NVFP4 serving)Mia X2026-09-05
~47 tok/s proseOne DGX SparkMia single-Spark vLLM NVFP4Mia X2026-09-06
24.5 tok/s batch-1 greedy, MTP off2× GB10, TP2+EP, 262K ctx, GMU 0.835vllm/vllm-openai:qwen38-flash-nextDual README (Mia-measured on their running container)accessed 2026-09-07
52.1 tok/s batch-1 greedy, MTP=3 (2.13× vs MTP off)Same two-Spark kitSame image; draft acceptance 72.8% in that tableSame README MTP tableaccessed 2026-09-07
54.4 tok/s prose ×1 (sparkDash)Same kitfp8 KV plus 65,536-id balanced MTP draft vocab — README says that vocab is not the defaultSame README Performance sectionaccessed 2026-09-07

Read the last row carefully. Mia’s README Performance table was captured with sparkDash against the running server at fp8 KV and a reduced MTP draft vocabulary passed via MTP_DRAFT_VOCAB. Drop that extra and the decode column moves; the README says so. The 2026-09-05 X post is the short public dual figure (54 tok/s prose single-stream). Do not collapse those rows into one LocalRig number.

The same Performance section lists aggregate decode rising to 207.0 tok/s at concurrency ×8 while per-stream falls to 26.7 tok/s, and prefill around 2.96k tok/s from 16k–64k prompts. Those are still Mia’s sparkDash captures on that kit. Cold start on the README’s default runtime is about 10m55s to API up.

None of these rows is a LocalRig measurement. Your interconnect, .env, image digest, and whether MTP drafting is on will not automatically reproduce them.

How do you follow Mia’s dual-Spark runbook without inventing steps?

Clone the public repo and run the scripts it ships; do not treat this article as a fork of the launch flags. The dual README Quick Start (accessed 2026-09-07) is four operator moves on the head node after you set IPs and the ConnectX interface in .env:

cp .env.sample .env
# edit HEAD_IP, WORKER_IP, IFACE, IB_HCA, and related fields in the repo's .env

./download.sh          # NVFP4 onto the head HF cache; ./download.sh --fp8 for official FP8
./start.sh --no-download
# confirm KV allocation once the server is up:
docker logs vllm-fn 2>&1 | grep -E "Available KV cache memory|GPU KV cache size"

That sequence is Mia’s, not LocalRig’s. Live flags, patches, and .env.sample values change; read the GitHub README on the day you launch. The same README documents --nfs to share the head Hugging Face cache instead of rsyncing ~126 GiB to the worker, ./start-fp8.sh for official FP8, and ./stop.sh before switching checkpoints. Both launch scripts put containers named vllm-fn on the same port.

What start.sh does, per that README: download (optional), distribute weights, sync the vLLM image to both nodes, bind-mount PLE and MXFP8 patches, refuse to launch if another process holds a GPU (REQUIRE_IDLE_GPU=true in sample), start the worker (rank 1) first, then the head (rank 0) on :8888. Shipped sample parallelism: --tensor-parallel-size 2, --nnodes 2, expert parallel, MTP 3. Default MAX_MODEL_LEN in .env.sample is 1,000,000 with YaRN; the README warns that YaRN at 1M was previously a silent no-op and is unvalidated on that kit after the Sep 2026 override fix — native 262,144 is the comparison length used in the measured KV and MTP tables.

Do not copy-paste a vllm serve … line from a blog cache. The README’s “Default runtime” block is a docker inspect dump from Mia’s live box (including that kit’s 10.0.0.1 master address). Your IPs, HCA names, and cross-wired interfaces will differ; .env.sample exists because of that.

If a step fails, the failure modes the README already names are the ones to check first: GPU not idle on either node, worker missing the ~126 GiB cache (unless NFS share is on), PLE table dtype / MTP expert patches not applied, MXFP8 shapes that need the bind-mounted fallback on GB10. Those patches are why this is a runbook, not a one-line vllm serve.

Should you buy two Sparks instead of dual RTX 3090s?

Buy two Sparks for this NVFP4 serving plan. Buy dual 3090s for 27B Q4. Do not use one purchase to stand in for the other. Two used RTX 3090s are 48 GB of split GDDR6X. That capacity class is the cheaper local path for Qwen 3.8 27B Q4. It is the wrong path for a 124 GiB-class NVFP4 checkpoint with a 51 GB PLE table.

Spark vs 3090 ownership (power, noise, unified vs discrete bandwidth) is already compared at DGX Spark vs dual RTX 3090. This page only adds the SKU split: Flash-Next NVFP4 → Spark runbooks; 27B Q4 → 24 GB cards. A 2-pack is the matching-systems bundle when you already know you want Mia’s TP2 recipe and the interconnect cable. A single Spark is the cheaper appliance experiment if you have not proven the software path yet.

DGX Spark 2-pack with cable → · Single DGX Spark →

When is a rented 80GB GPU the better try path?

When you want Unsloth’s 1-bit Flash GGUF (~75 GB), not when you want to rent a Spark. Vultr does not rent DGX Spark or GeForce 4090/3090 cards. If you do not want two Spark boxes and you are on the GGUF side of the table, an 80 GB A100 is a tight fit-check for Unsloth’s 1-bit band only — KV cache and MTP eat the margin; 4-bit in Unsloth’s table is 96–114 GB and does not fit one 80 GB GPU. Check current A100/H100 inventory on Vultr GPU Cloud. That rental is not a substitute for Mia’s dual-Spark NVFP4 interconnect.

For 24 GB GGUF experiments (27B, not this SKU), marketplace cards are the usual try path; they still do not emulate two GB10 nodes with ConnectX.

Who this is NOT for

  • Anyone who wants a LocalRig tok/s winner. There is no first-party dual-Spark row in engine.db. Every speed figure above is Mia’s dated public config.
  • Buyers whose target is Qwen 3.8 27B Q4. A used 3090 is the cheaper machine. This NVFP4 stack is the wrong spend for that SKU.
  • People who expect two Sparks to become 256 GB of one pool. TP2+EP serving is not a merged unified-memory computer.
  • Operators mixing Unsloth’s 75 GB GGUF with this vLLM NVFP4 recipe. Different checkpoint, different runtime, different buy list.
  • Dual-3090 builders hoping 48 GB split will load this Flash-Next NVFP4 checkpoint. Wrong capacity class.
  • Anyone treating vendor FP4 TOPS or a screenshot tok/s as a purchase guarantee. Fit, interconnect, and launch patches are separate from a single posted decode number.

Methodology

This article is a public-runbook how-to, not a LocalRig benchmark. No lab hour was run on DGX Spark for this page. Throughput, KV-cache sizes, cold-start time, and launch flags are copied from:

  1. MiaAI-Lab dual-Spark GitHub README, fetched 2026-09-07 from https://github.com/MiaAI-Lab/Qwen3.8-Flash-Next-Dual-DGX-Sparks (Quick Start, .env sample, KV budget, MTP table, sparkDash Performance section).
  2. Mia X posts dated 2026-09-05 (dual, 54 tok/s prose single-stream) and 2026-09-06 (single Spark, ~47 tok/s prose).
  3. The sibling single-Spark repo URL, used as the one-box path, not re-benched here.

Where the README marks a figure as measured on their running vllm-fn container, this page repeats that attribution. Where a Performance row includes a non-default MTP_DRAFT_VOCAB, that extra is named in the table. Commands in the how-to section are the public Quick Start from that README snapshot; they are not a LocalRig-maintained fork. Re-read the GitHub copy before launching — patches and defaults move.

Sources

Frequently Asked Questions

Do two DGX Sparks become one 256GB memory pool?

No. Each Spark still has its own 128GB unified-memory computer. Mia's dual runbook is two-node vLLM with tensor parallel size 2 plus expert parallel, over ConnectX. That shards the serving job; it does not merge the boxes into one 256GB machine.

Is Mia's dual-Spark NVFP4 the same as Unsloth's 75GB Flash GGUF?

No. Dual Spark here is vLLM NVFP4 on two GB10 appliances. Unsloth's Flash-Next GGUF is a 75GB-at-1-bit llama.cpp / Unsloth Desktop path sized for a 96GB RAM or unified box. Same family name, different checkpoint, runtime, and buy list.

Can two RTX 3090s run this Flash-Next SKU?

Not as this page's stack. Dual 3090s give 48GB split GDDR6X, which is the cheaper path for Qwen 3.8 27B Q4. Mia's NVFP4 checkpoint is on the order of 124 GiB on disk with a large PLE table; 48GB split is the wrong capacity class.

What public tok/s numbers exist for two Sparks?

Mia posted 54 tok/s prose single-stream on 2026-09-05 for the dual update. The dual GitHub README, accessed 2026-09-07, also lists Mia-measured sparkDash prose at 54.4 tok/s batch-1 on their two-GB10 kit with TP2+EP, MTP=3, 262K context. Those are Mia's numbers, not LocalRig lab hours.

Does LocalRig have a first-party dual-Spark benchmark?

No. This page cites dated public GitHub and X configs only. Do not treat any tok/s here as a LocalRig measurement or as a promise that your interconnect and .env will match Mia's kit.

Sources

  • MiaAI-Lab, Qwen3.8-Flash-Next-Dual-DGX-Sparks README, https://github.com/MiaAI-Lab/Qwen3.8-Flash-Next-Dual-DGX-Sparks, accessed 2026-09-07
  • MiaAI-Lab, Qwen3.8-Flash-Next-Single-DGX-Spark sibling runbook, https://github.com/MiaAI-Lab/Qwen3.8-Flash-Next-Single-DGX-Spark, accessed 2026-09-07
  • Mia (@MiaAI_lab), dual Spark 54 tok/s prose single-stream, 2026-09-05, https://x.com/MiaAI_lab/status/2096273630330081495
  • Mia (@MiaAI_lab), single Spark ~47 tok/s prose, 2026-09-06, https://x.com/MiaAI_lab/status/2096601527456592223
  • RadixArk/Qwen3.8-Flash-Next-NVFP4, https://huggingface.co/RadixArk/Qwen3.8-Flash-Next-NVFP4, accessed 2026-09-07
  • NVIDIA DGX Spark product page: 128GB unified memory, GB10, ConnectX-7 (nvidia.com, accessed 2026-07-29)