gpu4.pics
data via LocalScore
AMD Radeon RX 6650 XT
llamafile 0.9.1
8 GB VRAM AMD
running Meta Llama 3.1 8B Instruct· Q4_K_M
25.9
tokens / sec
181 · Modest
Generation
25.9tok/s
higher is better
Prompt / prefill
563tok/s
higher is better
Time to first token
2.48 s
lower is better
Composite score
181
Model size
8B
Runtime
llamafile
AMD EPYC 7352 24-Core Processor (znver2) · 126 GB RAM · Linux · llamafile
gpu4.pics/r/ls-result-13

All benchmarks on this rig

ModelQuantGenPromptTTFTPowerPerf/W$/MtokScore
Meta Llama 3.1 8B InstructQ4_K_M25.95632.48 s181

How it stacks up — Meta Llama 3.1 8B Instruct

AMD Radeon RX 6650 XT
25.9
NVIDIA H100
120
NVIDIA A100 80GB
110
NVIDIA RTX 3090 Ti
110
NVIDIA RTX PRO 6000 Blackwell Workstation Edition
103
NVIDIA RTX 3090
103
NVIDIA RTX A6000
90.5