gpu4.pics
data via LocalScore
NVIDIA RTX 5090
32 GB VRAM NVIDIA 575 W TDP
running Llama 3.2 1B Instruct· Q4_K_M· 1.5B
170
tokens / sec
3,245 · Excellent #15 of 768
Generation
170tok/s
higher is better
Prompt / prefill
20,305tok/s
higher is better
Time to first token
300 ms
lower is better
Composite score
3,245
Model size
1.5B
Runtime
gpu4.pics/r/ls-nvidia-rtx-5090-llama-32-1b-instruct-q4-k-m

All benchmarks on this rig

ModelQuantGenPromptTTFTPowerPerf/W$/MtokScore
Llama 3.2 1B Instruct1.5BQ4_K_M17020,305300 ms3,245

How it stacks up — Llama 3.2 1B Instruct

NVIDIA RTX 5090
170
NVIDIA L40S
355
NVIDIA RTX 3090 Ti
354
NVIDIA H100
335
NVIDIA RTX 3090
330
NVIDIA RTX A6000
315
NVIDIA A100 80GB
308