Will my GPU run it?

Approximate VRAM fit for each accelerator × model size, at a given quant and context. Estimated from weights + KV cache — the one question every local-AI builder asks, in one matrix.

✓ fits≈ tight✕ won't fit
AcceleratorVRAM1B3B8B14B32B70B120B180B
Apple M3 Ultra512 GB
Apple M2 Ultra192 GB
NVIDIA H200141 GB
Apple M3 Max128 GB
Apple M4 Max128 GB
AMD EPYC 7352 24-Core Processor (znver2)126 GB
NVIDIA RTX PRO 6000 Blackwell Server Edition95 GB
NVIDIA RTX PRO 6000 Blackwell Workstation Edition95 GB
NVIDIA A100 80GB80 GB
NVIDIA H10080 GB
NVIDIA RTX 6000 Ada Generation48 GB
NVIDIA RTX A600048 GB
NVIDIA L40S44 GB
AMD Instinct MI50 32GB32 GB
NVIDIA RTX 509032 GB
NVIDIA Tesla V10032 GB
NVIDIA RTX 309024 GB
NVIDIA RTX 3090 Ti24 GB
NVIDIA RTX 409024 GB
NVIDIA Tesla P4024 GB
Quadro RTX 600023 GB
NVIDIA GeForce RTX 4060 Ti16 GB
NVIDIA RTX 408016 GB
NVIDIA RTX 308010 GB
AMD Radeon RX 6650 XT8 GB
Apple M18 GB
NVIDIA GeForce RTX 2070 Super8 GB
NVIDIA GeForce RTX 3070 Ti8 GB

Estimates use effective bits-per-weight per quant + a KV-cache approximation; real fit varies with architecture, attention type, and runtime. Apple unified memory is shown as total RAM (usable VRAM is typically ~75%). Treat ≈/✕ near the boundary as “try it.”