gpu4.pics
data via LocalScore
NVIDIA RTX 3090
llamafile 0.9.2
24 GB VRAM NVIDIA 350 W TDP
running Llama 3.2 1B Instruct· Q4_K_M· 1.5B
308
tokens / sec
3,294 · Excellent
Generation
308tok/s
higher is better
Prompt / prefill
12,344tok/s
higher is better
Time to first token
106 ms
lower is better
Composite score
3,294
Model size
1.5B
Runtime
llamafile
Intel Xeon CPU E5-2660 v3 @ 2.60GHz (haswell) · 63 GB RAM · Linux · llamafile
gpu4.pics/r/ls-result-23

All benchmarks on this rig

ModelQuantGenPromptTTFTPowerPerf/W$/MtokScore
Llama 3.2 1B Instruct1.5BQ4_K_M30812,344106 ms3,294

How it stacks up — Llama 3.2 1B Instruct

NVIDIA RTX 3090
308
NVIDIA L40S
355
NVIDIA RTX 3090 Ti
354
NVIDIA H100
335
NVIDIA RTX 3090
330
NVIDIA RTX A6000
315
NVIDIA A100 80GB
308