Can I Run GLM-5.3-Flash Locally? 320B-A18B Memory Math, Not a 24 GB Card
Short answer: GLM-5.3-Flash is a 320B-A18B MoE (ox-alpha). Unsloth’s hardware table (accessed 2026-09-07) wants ~100 GB total memory at 1-bit and 128–150 GB at 3-bit. A used RTX 3090 or 4090 is the wrong purchase for this SKU. The honest local paths are a 128 GB Mac / DGX Spark at 3-bit, a large RAM box with offload, or a multi-GPU datacenter rental. LocalRig has not measured tok/s on this model.
This page is a constraint selector. It is not a quality ranking against Claude or GPT, and it is not the older GLM-5.2 flagship guide. Mixing those two names is how people buy a 24 GB card for a 100 GB weight file.
Is GLM-5.3-Flash the same model as GLM-5.2?
No. GLM-5.2 is Zhipu’s larger 1M-context flagship; the LocalRig 5.2 page still treats it as a datacenter / 256 GB-class problem. GLM-5.3-Flash is the smaller Flash SKU Unsloth documents as 320B parameters, 18B active, also called ox-alpha (Unsloth docs, GGUF on Hugging Face).
Unsloth’s own quality table compares Flash against GLM-5.2 and several closed models. That is a vendor scorecard, not a LocalRig ranking, and it is not a VRAM table. If you came here because an X thread stacked “GLM-5.3-Flash above Qwen 3.8 27B,” treat that as preference. The buy question is whether ~100 GB of weights plus KV cache fit the machine you actually own.
How much memory does GLM-5.3-Flash actually need?
About 100 GB at 1-bit, 128 GB-class at 3-bit, hundreds of GB at higher precision. Unsloth’s table is total memory: RAM + VRAM, or unified memory. Copied 2026-09-07:
| Quant band (Unsloth) | Total memory they list |
|---|---|
| 1-bit | 100 GB |
| 2-bit | 115 GB |
| 3-bit | 128–150 GB |
| 4-bit | 162–210 GB |
| 8-bit | 350 GB |
| BF16 | 650 GB |
File sizes on the same page (Dynamic GGUF, not LocalRig measurements): UD-IQ1_S 93.09 GB, UD-IQ1_M 97.58 GB, UD-Q2_K_XL 108.72 GB, UD-IQ3_XXS 120.37 GB, UD-Q3_K_XL 147.54 GB, UD-Q4_K_XL 199.71 GB. Unsloth’s 3-bit demo path is UD-IQ3_XXS because it is the quant they say fits 128 GB devices (Mac or NVIDIA DGX Spark).
Context is 1,048,576 tokens on the vendor card. Long context is extra KV cache on top of those weight files. A machine that “just fits” 1-bit at short context will not magically hold Max-thinking plus a huge window.
Use the VRAM calculator for a rough KV overlay, but treat Unsloth’s bands as vendor estimates. LocalRig has not loaded this GGUF.
Can a 24 GB card run GLM-5.3-Flash?
No. A 24 GB RTX 3090 or 4090 is about one-quarter of Unsloth’s 1-bit floor. Two used 3090s give 48 GB of aggregate device memory, which is still half of 1-bit and far under 3-bit. This is not a “Q4 it until it fits” problem. The 1-bit file is already ~93 GB.
If your silicon is a single 24 GB card, the current dense daily driver in this niche is Qwen 3.8 27B Q4, not this MoE. A used 3090 is still a valid buy for 27B. It is not a GLM-5.3-Flash buy.
Browse used RTX 3090 24GB on eBay → only if you are shopping 27B-class models. Do not buy a 3090 expecting ox-alpha to load.
What local machines actually fit Unsloth’s table?
128 GB unified memory is the 3-bit story. 96–128 GB RAM boxes are a 1-bit offload story, not a 24 GB story.
Unsloth states the smallest 1-bit quant works on 100 GB RAM and 3-bit on 128 GB devices like a Mac or NVIDIA DGX Spark. That maps to:
- Mac Studio / Mac Pro 128 GB (or more) unified memory — fit at 3-bit is the claim; decode speed is a different question. LocalRig still has no named Mac tok/s for this SKU.
- NVIDIA DGX Spark (128 GB unified) — same 3-bit fit class. Named Spark speed packets (llama.cpp 32K decode; vLLM short C1) are in the speed section below — they are not Mac numbers. Spark vs dual-3090 ownership: DGX Spark vs 2× RTX 3090. Spark does not turn two 24 GB cards into 128 GB shared.
- Workstation with ≥100 GB combined RAM+VRAM — possible for 1-bit with host offload. Speed then tracks the offload split, which LocalRig has not measured.
NVIDIA DGX Spark on Amazon → · Spark 2-pack →
Spark is an expensive way to buy 128 GB unified memory. It is a coherent path for this SKU’s 3-bit table. It is a waste if you only wanted 27B Q4.
How fast is it? Which numbers are public, and on what hardware?
The published llama-bench table is 1× B200, not a home GPU. Unsloth (Sep 4, 2026 note: “3.3× faster” plus MTP) reports UD-IQ1_S on 1× B200, MTP off first:
| Test | Baseline tok/s | Optimized tok/s |
|---|---|---|
| pp512 | 1121.80 | 1122.0 |
| tg32 | 62.79 | 63.10 |
| tg32 @ 4096 | 53.52 | 59.50 |
| tg32 @ 16384 | 41.02 | 57.99 |
| tg32 @ 65536 | 20.66 | 48.99 |
Those are B200 figures from Unsloth’s docs, accessed 2026-09-10. They are not RTX 3090, RTX 4090, Spark, or Mac numbers. With MTP, Unsloth shows further gains (example: prompt 4096, MTP off 58.6 vs n=2 86.5 tok/s) and warns that more draft tokens can slow the run. Copy the hardware label with the number or do not copy the number.
Two named Spark packets now exist. They are community-cited, not LocalRig-measured, and they are not an A/B: different quants and different prompt lengths.
| Source | Date | Hardware | Runtime | Quant | What was measured | tok/s |
|---|---|---|---|---|---|---|
| Weschera | 2026-08-27 | 1× DGX Spark GB10, 121 GiB UMA | llama.cpp (eauchs PR #27752, --spec-type draft-mtp) | Unsloth UD-Q2_K_XL (109 GB) | Decode at 32K context (temp 0, warmed). 262,144 boots with q8_0 KV; that is allocation, not this table. | Structured 33.8 · JSON 26.6 · code 25.0 · prose 17.9 |
| gitcommit90 | 2026-09-03 | 1× DGX Spark, 128 GB UMA | vLLM TP1 + DFlash2 K5 (default) | Turboderp EXL3 2.05 (85.23 GB) | Short C1 decode (400 out). Long-context column in that repo is prefill at 8k/16k/100k, not filled-262k decode. DFlash2 is CC BY-NC-ND 4.0. | Prose 29.9 · code 40.1 · structured 53.5 |
Do not average those rows. Do not paste them onto a Mac. Mac llama.cpp / Mac vLLM tok/s is still a shortfall — Unsloth names 128 GB Mac as a 3-bit fit class; llmcheck-style 32 tok/s figures are estimates, not a named packet.
LocalRig has no first-party GLM-5.3-Flash row in engine.db. Community X preference posts are not a substitute for a pinned config.
Should I rent an A100 or H100 to try it?
A single 80 GB A100 or H100 does not hold Unsloth’s 3-bit 128 GB band. 1-bit (~93–100 GB) also does not sit inside one 80 GB GPU without spilling into host RAM. Vultr’s catalog is datacenter GPUs (A100, H100, L40S), not consumer 4090s — useful if you need multi-GPU inventory or an 80 GB card for a different model that actually fits 80 GB (for example Unsloth’s Qwen 3.8 Flash GGUF 1-bit at ~75 GB). It is the wrong one-click assumption for GLM-5.3-Flash 3-bit.
Check current multi-GPU / 80 GB SKUs on Vultr GPU Cloud. Do not treat that link as “this card runs ox-alpha 3-bit.” For a 24 GB consumer-card experiment on a different model, use RunPod or Vast.ai. Full IaaS positioning: Vultr GPU cloud review.
How do I load it if the machine actually fits?
Unsloth documents Unsloth Desktop and a llama.cpp branch (glm5next/upstream on their fork). Their 3-bit demo command (vendor docs, accessed 2026-09-07):
hf download unsloth/GLM-5.3-Flash-GGUF \
--local-dir unsloth/GLM-5.3-Flash-GGUF \
--include "*UD-IQ3_XXS*"
./llama.cpp/llama-cli \
--model unsloth/GLM-5.3-Flash-GGUF/UD-IQ3_XXS/GLM-5.3-Flash-UD-IQ3_XXS-00001-of-00004.gguf \
--temp 1.0 \
--top-p 0.95 \
--chat-template-kwargs '{"reasoning_effort":"max"}'
Default sampling in that guide is temperature=1.0, top_p=0.95. Reasoning effort is low / high / max. That is vendor procedure, not a LocalRig lab notebook. Build the runtime they name; old llama.cpp tags will fail.
Related: Ollama vs llama.cpp vs vLLM.
No new Spark/Mac × llama.cpp vs vLLM page. Spark now has two named packets (llama.cpp GGUF vs vLLM EXL3) on this selector. Mac still has no named tok/s. The quants differ, and neither Spark row is filled-200k decode. An Ollama vs llama.cpp vs vLLM triangle would still be spam. The B200 llama-bench stays labeled B200. The 24 GB GLM Flash in this batch is GLM-4.7-Flash (30B-A3B, Unsloth ~18 GB 4-bit) — not ox-alpha.
Who this is NOT for
- Anyone who wants LocalRig tok/s on GLM-5.3-Flash. There is no first-party row. The B200 table is Unsloth’s. The Spark rows are community GitHub recipes, not a Mac number and not filled-200k.
- Buyers with a 12–24 GB Ampere or Ada card. Wrong SKU. Use Qwen 3.8 27B.
- People collapsing this into GLM-5.2. Different model, different memory class. Keep the 5.2 page for the flagship.
- Anyone assuming one 80 GB A100 is a 3-bit fit. 128–150 GB is not 80 GB.
- Shoppers treating vendor “rivals Opus” tables as a hardware requirement. Quality scorecards do not change the 93 GB file.
Methodology
- Fit numbers: Unsloth hardware table and Dynamic GGUF file sizes, accessed 2026-09-07 from https://unsloth.ai/docs/models/glm-5.3-flash and the Hugging Face GGUF repo. Labeled vendor estimates.
- Speed numbers: Unsloth llama-bench on 1× B200, UD-IQ1_S, with and without their Sep 4 optimizations / MTP. Spark: Weschera llama.cpp
UD-Q2_K_XLat 32K (2026-08-27); gitcommit90 vLLM EXL3 2.05 DFlash2 short C1 (2026-09-03). Not LocalRig-measured. Not transferable to 3090 / 4090 / Mac. Not filled-200k decode. - SKU identity: Unsloth docs (320B-A18B, ox-alpha) plus the existing LocalRig GLM-5.2 selector for the name split.
- LocalRig first-party: none for GLM-5.3-Flash as of 2026-09-10. No
engine.dbbenchmark row; this page is not a model×hardware matrix article.
Sources
- Unsloth GLM-5.3-Flash local guide — memory table, 1-bit/3-bit floors, B200 llama-bench, MTP note (Sep 4, 2026). Accessed 2026-09-10.
- unsloth/GLM-5.3-Flash-GGUF — GGUF files.
- Weschera/glm53-flash-one-spark — 1× Spark llama.cpp,
UD-Q2_K_XL, 32K decode, 2026-08-27. - gitcommit90/glm-5.3-one-spark — 1× Spark vLLM EXL3 2.05 + DFlash2, 2026-09-03.
- Can I run GLM-5.2 locally? — different SKU.
- Can I run GLM-4.7-Flash locally? — 24 GB GLM Flash, not this 100 GB SKU.
- Can I run Qwen 3.8 locally? — 24 GB dense path.
- DGX Spark vs 2× RTX 3090 — 128 GB unified vs 48 GB split.
- Vultr GPU Cloud — datacenter inventory; not a consumer 4090 and not a promised 3-bit single-GPU fit.
Frequently Asked Questions
Can I run GLM-5.3-Flash on a used RTX 3090 or 4090?
No. Unsloth's 1-bit band is about 100 GB of total RAM+VRAM. A 24 GB card cannot hold the weights. Two 3090s (48 GB aggregate) still miss the floor. This is a 128 GB-class machine, a multi-GPU datacenter box, or a skip.
Is GLM-5.3-Flash the same as GLM-5.2?
No. GLM-5.2 is the larger 1M-context flagship that still does not fit home cards. GLM-5.3-Flash is ox-alpha, 320B total / 18B active. Do not reuse the 5.2 math on this SKU.
Will an 80 GB A100 or H100 run Unsloth's 3-bit GGUF?
A single 80 GB card does not hold Unsloth's 3-bit band (128–150 GB). 1-bit (~100 GB file ~93 GB) also overflows one 80 GB GPU unless you offload heavily into host RAM. Check a 128 GB Mac / Spark, or a multi-GPU datacenter inventory — not a one-card A100 assumption.
How fast is GLM-5.3-Flash locally?
LocalRig has no first-party row. Unsloth's llama-bench for UD-IQ1_S is 1× B200 (tg32 about 63 tok/s). Two named Spark packets exist: Weschera llama.cpp UD-Q2_K_XL at 32k context (17.9–33.8 tok/s, 2026-08-27) and gitcommit90 vLLM EXL3 2.05 DFlash2 at short C1 (29.9 prose / 40.1 code, 2026-09-03). Neither is filled-200k decode. No named Mac tok/s.
What should I buy instead if I only have 24 GB?
Run Qwen 3.8 27B Q4 on the 24 GB card. GLM-5.3-Flash is the wrong download for that machine. See the Qwen 3.8 27B selector.