gpu4.pics
data via LocalScore
NVIDIA RTX PRO 6000 Blackwell Workstation Edition
95 GB VRAM NVIDIA
running Llama 3.2 1B Instruct· Q4_K_M· 1.5B
215
tokens / sec
3,948 · Excellent #6 of 768
Generation
215tok/s
higher is better
Prompt / prefill
22,640tok/s
higher is better
Time to first token
125 ms
lower is better
Composite score
3,948
Model size
1.5B
Runtime
gpu4.pics/r/ls-nvidia-rtx-pro-6000-blackwell-workstation-edition-llama-32-1b-instruct-q4-k-m

All benchmarks on this rig

ModelQuantGenPromptTTFTPowerPerf/W$/MtokScore
Llama 3.2 1B Instruct1.5BQ4_K_M21522,640125 ms3,948

How it stacks up — Llama 3.2 1B Instruct

NVIDIA RTX PRO 6000 Blackwell Workstation Edition
215
NVIDIA L40S
355
NVIDIA RTX 3090 Ti
354
NVIDIA H100
335
NVIDIA RTX 3090
330
NVIDIA RTX A6000
315
NVIDIA A100 80GB
308