Can my GPU run this model?
Enter the model's parameter count, quantization level, and your target context window. The calculator estimates required VRAM (weights + KV cache) and shows which hardware tiers comfortably fit, with links to buy.
Calculation is approximate. Actual VRAM usage varies by runtime, batch size, and system reserved memory. Results are conservative by design — they include a 15% headroom buffer.
Inputs
Parameter count × quant × context. Default is 200k — the window long-running agents and coding work actually use. KV cache, not weights, dominates from here up to 1M.
Set a model and hit Run.
Common hardware VRAM reference
Sorted by VRAM tier. Prices are approximate used-market estimates as of mid-2026 — verify before buying.
| GPU / Hardware | VRAM | BW (GB/s) | Price est. | Best for | Links |
|---|---|---|---|---|---|
| RTX 4060 Ti 8GB | 8 GB | 288 | $280–$370 used | 7B Q4 only | Amazon · eBay used |
| RTX 4060 Ti 16GB | 16 GB | 288 | $400–$480 used | 7B Q8, 13B Q4 (slow BW) | Amazon · eBay used |
| RTX 4070 Ti Super 16GB | 16 GB | 672 | $650–$800 used | 7B Q8, 13B Q4 | Amazon · eBay used |
| AMD RX 7900 XT 20GB | 20 GB | 800 | $450–$600 used | 13B Q8, 34B Q2 (ROCm) | eBay used |
| RTX 3090 24GB ⭑ | 24 GB | 936 | $550–$750 used | 7B–34B Q4, best $/VRAM | eBay used |
| RTX 4090 24GB | 24 GB | 1008 | $1,350–$1,600 used | 7B–34B, fastest consumer | Amazon · eBay used |
| AMD RX 7900 XTX 24GB | 24 GB | 960 | $650–$850 used | 7B–34B Q4 (ROCm) | Amazon · eBay used |
| Mac Mini M4 Pro 48GB | 48 GB | 273 | ~$1,799 new | 34B Q4–Q8, silent, efficient | Amazon |
| NVIDIA A6000 48GB ⭑ | 48 GB | 768 | $2,000–$3,000 used | 70B Q4 in VRAM, best $/48GB | eBay used |
| RTX 6000 Ada 48GB | 48 GB | 960 | $4,500–$6,000 used | 70B Q4–Q6, fastest 48GB | eBay used |
| Mac Studio M4 Max 128GB | 128 GB | 546 | ~$3,999 new | 70B Q4–Q8, silent | Amazon |
| Mac Studio M4 Ultra 192GB | 192 GB | 1092 | ~$4,999 new | 70B Q8, 100B+ class, long context | Amazon |
| NVIDIA DGX Spark 128GB | 128 GB | — | appliance kit | 200k-class context on 27B–70B; not a 24 GB substitute | Spark vs dual-3090 |
| Dual DGX Spark (two-node) | 256 GB pool | — | two units | 1M-class windows that a single 128 GB box cannot hold | Dual-Spark how-to |
⭑ = top value pick at that tier. Prices verified mid-2026; VRAM prices fluctuate with GPU releases. Always check current sold listings before purchasing.
Links are affiliate links (Amazon Associates · eBay Partner Network). Commission is earned at no extra cost to you. Picks are based on benchmark data, not commission rates.
Quick answers
How much VRAM do I need for Qwen 3.8 27B?
About 17–24 GB at Q4_K_M with a short session (4k–32k). That is a used RTX 3090 decision for short prompts. Agent and coding context (128k–200k) needs 128 GB unified memory or a cluster — the 24 GB card will OOM on KV cache.
Can a 12 GB GPU run an 8B model?
Yes at Q4_K_M with a short context. A used RTX 3060 12GB is the entry pick. Do not buy 12 GB if Qwen 3.8 27B is the goal.
Why does 70B Q4 want 48 GB?
Weights are ~40 GB plus KV cache. No current one-card consumer GPU holds that. Dual used 3090s or a used A6000 are the 48 GB floor; a 32 GB 5090 is still short.
Can I run 200k or 1M context locally?
Yes on enough memory — not on a 24 GB card. Local buyers here are running long agents and coding sessions, so 200k is the default window. A 27B Q4 that fits a used 3090 at 4k will not hold 200k, let alone 1M, without 128GB+ unified memory (Mac Studio, DGX Spark), a multi-GPU pool, or a two-node kit.
Get notified when new benchmarks drop
We publish hardware benchmarks and buying alerts when major models release. Weekly digest, no fluff.