gpu4.pics
data via LocalScore
Quadro RTX 6000
23 GB VRAM NVIDIA
running Llama 3.2 1B Instruct· Q4_K_M· 1.5B
253
tokens / sec
2,526 · Excellent #41 of 768
Generation
253tok/s
higher is better
Prompt / prefill
9,138tok/s
higher is better
Time to first token
143 ms
lower is better
Composite score
2,526
Model size
1.5B
Runtime
gpu4.pics/r/ls-quadro-rtx-6000-llama-32-1b-instruct-q4-k-m

All benchmarks on this rig

ModelQuantGenPromptTTFTPowerPerf/W$/MtokScore
Llama 3.2 1B Instruct1.5BQ4_K_M2539,138143 ms2,526

How it stacks up — Llama 3.2 1B Instruct

Quadro RTX 6000
253
NVIDIA L40S
355
NVIDIA RTX 3090 Ti
354
NVIDIA H100
335
NVIDIA RTX 3090
330
NVIDIA RTX A6000
315
NVIDIA A100 80GB
308