gpu4.pics
data via LocalScore
AMD Radeon RX 6650 XT
llamafile 0.9.2
8 GB VRAM AMD
running Llama 3.2 1B Instruct· Q4_K_M· 1.5B
82.9
tokens / sec
703 · Strong
Generation
82.9tok/s
higher is better
Prompt / prefill
2,592tok/s
higher is better
Time to first token
620 ms
lower is better
Composite score
703
Model size
1.5B
Runtime
llamafile
AMD EPYC 7352 24-Core Processor (znver2) · 128 GB RAM · Windows · llamafile
gpu4.pics/r/ls-result-19

All benchmarks on this rig

ModelQuantGenPromptTTFTPowerPerf/W$/MtokScore
Llama 3.2 1B Instruct1.5BQ4_K_M82.92,592620 ms703

How it stacks up — Llama 3.2 1B Instruct

AMD Radeon RX 6650 XT
82.9
NVIDIA L40S
355
NVIDIA RTX 3090 Ti
354
NVIDIA H100
335
NVIDIA RTX 3090
330
NVIDIA RTX A6000
315
NVIDIA A100 80GB
308