Benchmarks

These guides carry dated performance data — the hardware, runtime version, quantization, and the date each figure was collected. First-party results were measured by LocalRig; community-sourced figures inside a guide are labeled as such. How this testing works is on the methodology page.

Comparable benchmark records

Rows appear only when the hardware, model, runtime, test configuration, provenance, data date, methodology, and at least one measured result are recorded.

Filter complete records
Complete, attributed inference runs. Use the linked guide and methodology to inspect each result.
Hardware Model / quant Runtime / version Context / batch tok/s TTFT Memory Origin Date Methodology
Apple M4 16GB Mac mini Llama 3.1 8B Instruct Q4_K_M Hardware to Run a 7B/8B Model Locally: RTX 3090, Apple M3 Max, and Budget Options llama.cpp b9820 4,096 / 512 18.4 tok/s -- 8.0 GB first-party Jun 27, 2026 Methodology
Apple M4 16GB Mac mini Llama 3.1 8B Instruct Q4_K_M Hardware to Run a 7B/8B Model Locally: RTX 3090, Apple M3 Max, and Budget Options Ollama 0.30.11 4,096 / 512 19.5 tok/s -- 8.0 GB first-party Jun 27, 2026 Methodology

Benchmark guides with methodology

Dated guides remain available here when they document a method but do not provide a complete structured record suitable for a side-by-side comparison.

  • Hardware to Run a 70B Model Locally: VRAM, the 48GB Wall, and Your Real Options
    Data: Jul 10, 2026 What it actually takes to run a 70B model at home: the VRAM math, why 48GB is the practical floor at Q4, and the five hardware paths (dual 3090, used A6000, Apple Silicon, DGX Spark, or cloud).
  • The Local-AI Hardware Buying Framework
    Data: Jul 10, 2026 A constraint-first framework for choosing hardware to run AI models locally. Covers VRAM, memory bandwidth, quantization, Apple Silicon, DGX Spark, and budget paths — so you buy once and regret nothing.
  • Cudo Compute Review 2026: Distributed GPU Cloud Without the Marketplace Roulette
    Data: Jul 9, 2026 Cudo Compute occupies the middle ground between Vast.ai's bargain chaos and hyperscaler lock-in: distributed GPU supply with SMB-grade account management. Honest assessment of who it fits, pricing anchored to H100 market rates, and when RunPod or Vast is the better choice.
  • DigitalOcean GPU Droplets Review 2026: Beginner-Friendly, Not the Cheapest
    Data: Jul 9, 2026 DigitalOcean's GPU Droplets are the beginner-friendliest tier-1 cloud GPU option: predictable per-GPU-hour pricing, no spot preemption, and tight integration with existing DO infrastructure. Paperspace (acquired 2023) still runs as a separate subscription-gated product, not a merged one. For hobbyists escaping ChatGPT costs.
  • GPUMart Review 2026: Bare-Metal GPU Servers and the Real Rent-a-4090-Monthly Math
    Data: Jul 9, 2026 GPUMart's monthly bare-metal RTX 4090 rental creates one of the few apples-to-apples rent-vs-buy comparisons in the market. When high utilization and no upfront capital meet a sub-one-year horizon, monthly bare metal can beat both hourly cloud and ownership. This guide cuts through the pricing layers.
  • Lambda Cloud Review 2026: The $4.29 H100 Standard-Bearer for Serious Training
    Data: Jul 9, 2026 Lambda Cloud is the reference point builders compare against: clean datacenter H100s at ~$4.29/hr on-demand (single GPU) or ~$4.09/hr per GPU in an 8x cluster, with no marketplace variance. Wins decisively for multi-hour training and fine-tuning where interruption risk is high. Overkill for inference workloads where cheaper marketplaces undercut, and pricier than it used to be.
  • Modal Review 2026: 12-Second Cold Starts, and the Pricing Multipliers Nobody Reads
    Data: Jul 9, 2026 Modal's GPU snapshotting cut serverless cold starts from 118 seconds to 12 seconds. But non-preemptible, non-default-region pricing can quietly stack to 5x+ list cost. When serverless GPU makes sense, and when a rented pod wins.
  • RunPod Review 2026: Secure Cloud vs Community Cloud, Pricing, and the Data-Loss Gotcha
    Data: Jul 9, 2026 RunPod is the only tier-1 provider renting a real consumer RTX 4090. This guide covers Secure Cloud vs Community Cloud, when interruptions cost more than savings, the network-volume data-loss failure mode, and whether RunPod makes sense for fine-tuning or inference.
  • Salad Cloud Review 2026: $0.20/hr 4090s on Gaming PCs — Too Cheap to Be True?
    Data: Jul 9, 2026 Salad rents compute time on 60,000+ consumer gaming PCs at $0.20/hr for an RTX 4090. The catch is reliability: you get stateless, checkpointed, interruption-tolerant inference at a massive price advantage, or long stateful jobs fail mid-way. This guide shows when Salad wins and when to rent from a datacenter instead.
  • Is Vast.ai Safe? An Honest 2026 Review of the Cheapest GPU Marketplace
    Data: Jul 9, 2026 Vast.ai is legitimate, not a scam. But it is a peer-to-peer GPU marketplace with a variance problem: host reputations vary widely, pricing can creep past advertised rates, and unattended workloads silently fail. This guide separates real risks from rumors and tells you when Vast makes sense.
  • Vultr GPU Cloud Review 2026: Established IaaS for AI Workloads — Who It's Actually For
    Data: Jul 9, 2026 Vultr occupies the middle ground between hyperscaler pricing and GPU marketplace volatility. Datacenter-grade GPUs (A100, H100), hourly billing, no consumer chips, and a familiar control plane. The guide: when Vultr wins, why reliability costs, and who should look elsewhere.
  • Can I Run GLM-5.2 Locally? Hardware Requirements for the 1M-Context Flagship
    Data: Jun 29, 2026 GLM-5.2 (released June 2026) is open-weight but not runnable on most home hardware. The 1M-context model demands a 256GB-class machine and hits single-digit tokens/sec even then. Honest constraints, the math, and where to go instead.
  • AMD 7900 XTX for Local LLMs in 2026: ROCm Finally Grew Up
    Data: Jun 29, 2026 ROCm 7.2 (March 2026) closed the compatibility gap that used to make AMD a hard pass for local LLM inference. This guide covers what "parity" actually means for the 7900 XTX's 24GB, where it beats the RTX 4090 on price, where it loses on speed, and who should still buy CUDA.
  • Best GPU Under $500 for Local LLMs in 2026, Ranked by Constraint
    Data: Jun 29, 2026 At sub-$500, the GPU decision shifts from "what will fit" to "what will actually run without exploding the power supply." A constraint-ranked guide to the used 3060 12GB (safest default), Tesla P40 (VRAM-dense but cooling-complex), and the RTX 5060 Ti 16GB (best new-card path, just over budget).
  • GPU VRAM Tiers for Local AI: What 8, 12, 16, 24, 48GB Actually Buy You
    Data: Jun 29, 2026 A constraint-first selector that maps VRAM tiers—8GB, 12GB, 16GB, 24GB, 48GB—to the model classes they serve at usable quantization, with the heuristic to size your tier and the best-value GPU card per tier. No guesswork: know exactly which models fit where.
  • Is the RTX 5090 Worth It for Local AI in 2026? $2,000 MSRP, $3,700+ Reality
    Data: Jun 29, 2026 The RTX 5090 lists at $1,999 but is selling near double that in mid-2026. This guide runs the actual VRAM-per-dollar math against used RTX 3090 pairs and Apple Silicon unified memory, plus the PSU and PCIe 5.0 platform costs nobody puts on the spec sheet.
  • Two Used RTX 3090s vs One RTX 4090: The 48GB Question
    Data: Jun 29, 2026 For roughly the same spend, two used RTX 3090s buy 48GB of VRAM while one RTX 4090 buys 24GB at higher per-card speed. The real decision isn't value — it's whether your target model needs more than 24GB, because that single fact locks in a PSU, a motherboard, and a runtime you can't undo cheaply.
  • Why GPU Prices Spiked Again in 2026: The DRAM Shortage, Explained
    Data: Jun 29, 2026 2026's GPU price spike is rooted in DRAM memory shortage, not a repeat of the 2021–22 crypto boom. AI datacenters are locking up supply years ahead; the timeline to clearance is uncertain. Your honest options: buy used, switch to AMD, rent instead, or wait—and the right choice depends on your workload and timeline, not scarcity pressure.
  • M5 Ultra Mac Studio: Wait for It, or Buy M3 Ultra/M4 Max Now?
    Data: Jun 29, 2026 M5 Ultra is rumored for Q4 2026 with possible neural accelerator gains on prefill speed. But if you need a Mac now, know what actually improves from M3 Ultra to M4 Max, and when the wait genuinely pays for your workload.
  • Mac Studio Clusters over Thunderbolt 5: Trillion-Parameter Models at Home?
    Data: Jun 29, 2026 Four stacked Mac Studios over Thunderbolt 5, running macOS Tahoe RDMA, reportedly unlock trillion-parameter inference at home. A frontier-hobbyist path with honest cost-per-token math, attribution of unverified claims, and clarity on when cloud rental wins.
  • Mac Studio vs RTX 5090 for Local AI: Unified Memory vs Raw Speed
    Data: Jun 29, 2026 A widely-shared 2026 thread claims the RTX 5090 "can't keep up" with Apple Silicon — and on giant MoE models, it's true: 32GB of VRAM simply can't hold what 128-256GB of unified memory can. But the 5090 crushes anything that fits in its VRAM and owns fine-tuning outright. The real answer is a decision tree by model size, not a winner.
  • 10GbE vs 2.5GbE for an AI Homelab: When Model Loading Justifies the Upgrade
    Data: Jun 29, 2026 The math on pulling a 40GB 70B-Q4 GGUF from NAS: ~2.6–2.8 minutes at 2.5GbE versus ~42–43 seconds at 10GbE. A constraint-first decision tree: local NVMe first, network second, and why used enterprise SFP+ gear is the honest value path.
  • Beelink GTR9 Pro for Local LLMs: 128GB Unified Memory Under $2,000, Examined
    Data: Jun 29, 2026 The GTR9 Pro puts 128GB of unified memory in a silent mini-PC at $1,899. What does "70B Q5 ready" mean at this bandwidth tier? Honest spec breakdown, comparison to Mac Studio and discrete-GPU builds, and Beelink's support track record for the 70B-curious buyer.
  • Tesla V100 Budget AI Homelab: Datacenter Cast-Offs as the Value Multi-GPU Path
    Data: Jun 29, 2026 2× Tesla V100 32GB with HBM2 bandwidth and NVLink capability at half the cost of new consumer cards — but earned only with cooling mods, blower management, and SXM2 adapter tinkering. For builders, not plug-and-play buyers.
  • Cheapest RTX 5090 Cloud Rental in 2026: Prices, Availability, and Why Supply Is the Story
    Data: Jun 29, 2026 The RTX 5090 rental market is defined by supply constraints, not competition. A 7.4× price spread ($0.27–$2.00/hr) across providers reflects which ones have cards at all. Compare Salad, RunPod, CloudRift, and Vast.ai; break-even math against the ~$3,800 street price for ownership; and when renting beats buying for realistic workloads.
  • TensorDock Review 2026: The $0.37/hr RTX 4090 Marketplace, Honestly Assessed
    Data: Jun 29, 2026 TensorDock's spot RTX 4090 rentals at $0.20–$0.37/hr are among the cheapest on the market. But they're marketplace listings with no quality guarantees — identical constraint logic to Vast.ai. When they work, the unit economics are real. When they don't, you have no recourse. Honest pricing, zero affiliation, pure editorial trust piece.
  • Best GPU for Local LLM Inference (2026): VRAM-per-Dollar Guide
    Data: Jun 28, 2026 The GPU decision for local LLM inference is set by VRAM (does the model fit) and memory bandwidth (how fast it decodes), not raw FLOPS. A constraint-first, VRAM-per-dollar guide: used RTX 3090 vs RTX 4090 vs RTX 3060, multi-GPU reality, and when to switch to Apple Silicon.
  • Best Mac for Local LLM Inference (2026): The Unified-Memory Buying Guide
    Data: Jun 28, 2026 Two specs decide a Mac for local AI — unified memory (what fits) and memory bandwidth (how fast it decodes). The M-series ladder from Mac mini M4 to Mac Studio M3 Ultra, with first-party and community numbers and honest tradeoffs vs a discrete GPU.
  • How to Run LLMs Locally: Which Inference Engine for Your Rig (2026)
    Data: Jun 28, 2026 A decision guide that picks the right local inference engine from your hardware, not hype. llama.cpp for CPU and portability, MLX on Apple Silicon, vLLM for CUDA serving — and why we don't recommend Ollama.
  • How to Run llama.cpp on an RTX 3090 (CUDA, Step by Step)
    Data: Jun 28, 2026 A step-by-step guide to building llama.cpp with CUDA and running a GGUF model on an RTX 3090. Covers driver and toolkit prerequisites, the CUDA build, full GPU offload with -ngl, a throughput check, and an OpenAI-compatible server.
  • Hardware to Run a 7B/8B Model Locally: RTX 3090, Apple M3 Max, and Budget Options
    Data: Jun 27, 2026 Benchmark-backed hardware guide for running 7B and 8B parameter models locally. Covers RTX 3090, Apple M3 Max, RTX 3060, and Apple M4 — with first-party Apple M4 benchmarks, community throughput data, VRAM requirements, and honest trade-offs.
  • Quantization: What It Means for Local AI and Why It Matters
    Data: Jun 27, 2026 Quantization reduces the numerical precision of a model's weights to shrink its memory footprint — the single technique that determines whether a 7B or 70B model fits in your GPU's VRAM and how fast it will run.