gpu4.pics
data via LocalScore
NVIDIA RTX 6000 Ada Generation
48 GB VRAM NVIDIA
running Llama 3.2 1B Instruct· Q4_K_M· 1.5B
131
tokens / sec
3,350 · Excellent #13 of 768
Generation
131tok/s
higher is better
Prompt / prefill
19,620tok/s
higher is better
Time to first token
68 ms
lower is better
Composite score
3,350
Model size
1.5B
Runtime
gpu4.pics/r/ls-nvidia-rtx-6000-ada-generation-llama-32-1b-instruct-q4-k-m

All benchmarks on this rig

ModelQuantGenPromptTTFTPowerPerf/W$/MtokScore
Llama 3.2 1B Instruct1.5BQ4_K_M13119,62068 ms3,350

How it stacks up — Llama 3.2 1B Instruct

NVIDIA RTX 6000 Ada Generation
131
NVIDIA L40S
355
NVIDIA RTX 3090 Ti
354
NVIDIA H100
335
NVIDIA RTX 3090
330
NVIDIA RTX A6000
315
NVIDIA A100 80GB
308