GPUs
- Why VRAM (and Memory Bandwidth) Matters More Than Compute for Local LLMs
The single concept that explains the used-GPU market: LLM inference—especially token generation—is memory-bandwidth-bound, not compute-bound. Older cards with fat VRAM beat newer cards with thin VRAM. Learn capacity vs. bandwidth vs. compute, and why a 2020 Tesla P40 or 2021 AMD MI50 often outperform newer GPUs for local inference on a budget.
- Why GPU Prices Spiked Again in 2026: The DRAM Shortage, Explained
2026's GPU price spike is rooted in DRAM memory shortage, not a repeat of the 2021–22 crypto boom. AI datacenters are locking up supply years ahead; the timeline to clearance is uncertain. Your honest options: buy used, switch to AMD, rent instead, or wait—and the right choice depends on your workload and timeline, not scarcity pressure.
- Used RTX A6000 48GB for Local AI: The Single-Card Ceiling for 32B Models
As enterprises refresh to Blackwell, used RTX A6000 prices are softening. The A6000's 48GB on a single 300W card beats the multi-GPU wiring complexity of dual 3090s — but only if you accept the trade-offs: workstation cooling, no display output on some variants, and lower VRAM-per-dollar than two used 3090s combined.
- Tesla P40 in 2026: 24GB of VRAM for $150 — Legit Bargain or Trap?
The Tesla P40 costs $150–200 used and holds 24GB of VRAM. It runs local LLMs, but passive cooling throttles hard without forced airflow, compute is ancient, and there is no video output. A constraint guide: right for the tinkering homelabber who enjoys the project; wrong for anyone who wants inference to work on day one.
- RTX 4090 for Local LLMs in 2026: Great Card, Broken Price
The RTX 4090 is the fastest single-card GPU for LLM inference—~20% faster than an RTX 3090. But at $2,000+ used, it costs 2.5× as much for identical 24GB VRAM. This guide covers when the premium makes sense (prompt speed, diffusion, multi-user serving) and when you should buy a 3090 or rent instead.
- RTX 4060 Ti 16GB vs 5060 Ti 16GB: The Budget Pick That Changed in Months
The 4060 Ti 16GB was the default sub-$500 local LLM recommendation for two years. The 5060 Ti 16GB arrived near the same price with meaningfully better memory bandwidth, and bandwidth — not compute — is what token decode actually spends. Here is the price-conditional verdict, not a blanket "buy the new one."
- RTX 3060 12GB for Local LLMs: The Honest Entry Point Under $300
A 12GB VRAM card for under $300 can run 7B–8B models comfortably at Q4 quantization. This is the right first GPU when budget rules everything; the wrong choice if you already know you want bigger models. We show the model-fits math, the speed ceiling, and the upgrade path.
- Is NVLink Worth It in 2026? The Dual-3090 Bridge Question, Answered
NVLink is a $40–80 bridge with outsized mythology. For inference workloads that layer-split models across cards, it barely matters. For tensor-parallel serving it can. This guide separates the workloads where the bridge pays for itself from those where it's cargo cult tech.
- How to Buy a Used GPU Without Getting Burned: Vetting, Testing, and Red Flags
The used GPU market holds real value — and real risk. This guide walks through seller red flags, the first-48-hours test protocol (VRAM testing, sustained load, thermals), and return-window discipline. Persona 2's defective VRAM module risk is why this page exists.
- GPU VRAM Tiers for Local AI: What 8, 12, 16, 24, 48GB Actually Buy You
A constraint-first selector that maps VRAM tiers—8GB, 12GB, 16GB, 24GB, 48GB—to the model classes they serve at usable quantization, with the heuristic to size your tier and the best-value GPU card per tier. No guesswork: know exactly which models fit where.
- Best GPU Under $500 for Local LLMs in 2026, Ranked by Constraint
At sub-$500, the GPU decision shifts from "what will fit" to "what will actually run without exploding the power supply." A constraint-ranked guide to the used 3060 12GB (safest default), Tesla P40 (VRAM-dense but cooling-complex), and the RTX 5060 Ti 16GB (best new-card path, just over budget).
- AMD MI50 32GB: The $150 HBM2 Wildcard for Local LLMs
The MI50 offers ~2.5× memory bandwidth of an RTX 3090 at 10-15% of the cost. The catch is equal in size: ROCm support is in maintenance mode, so drivers will not improve and may decay. A knowing bet for tinkerers; explicitly not a recommendation for anyone's only GPU.
- AMD 7900 XTX for Local LLMs in 2026: ROCm Finally Grew Up
ROCm 7.2 (March 2026) closed the compatibility gap that used to make AMD a hard pass for local LLM inference. This guide covers what "parity" actually means for the 7900 XTX's 24GB, where it beats the RTX 4090 on price, where it loses on speed, and who should still buy CUDA.
- Is the RTX 5090 Worth It for Local AI in 2026? $2,000 MSRP, $3,700+ Reality
The RTX 5090 lists at $1,999 but is selling near double that in mid-2026. This guide runs the actual VRAM-per-dollar math against used RTX 3090 pairs and Apple Silicon unified memory, plus the PSU and PCIe 5.0 platform costs nobody puts on the spec sheet.
- Two Used RTX 3090s vs One RTX 4090: The 48GB Question
For roughly the same spend, two used RTX 3090s buy 48GB of VRAM while one RTX 4090 buys 24GB at higher per-card speed. The real decision isn't value — it's whether your target model needs more than 24GB, because that single fact locks in a PSU, a motherboard, and a runtime you can't undo cheaply.
- Used RTX 3090 Buying Guide 2026: Still the Best $/VRAM in Local AI — If You Vet It Right
The used RTX 3090 remains the consensus VRAM-per-dollar champion for local LLM inference in mid-2026, but the used market carries real betrayal risk: defective VRAM modules, undisclosed mining-farm history, and PSUs sized for gaming instead of sustained AI load. This is the vetting checklist and the honest "when not to buy one" case.
- Best GPU for Local LLM Inference (2026): VRAM-per-Dollar Guide
The GPU decision for local LLM inference is set by VRAM (does the model fit) and memory bandwidth (how fast it decodes), not raw FLOPS. A constraint-first, VRAM-per-dollar guide: used RTX 3090 vs RTX 4090 vs RTX 3060, multi-GPU reality, and when to switch to Apple Silicon.