gpu4.pics
data via LocalScore
NVIDIA RTX 3090
llamafile 0.9.1
24 GB VRAM NVIDIA 350 W TDP
running Meta Llama 3.1 8B Instruct· Q4_K_M
103
tokens / sec
1,008 · Excellent
Generation
103tok/s
higher is better
Prompt / prefill
3,552tok/s
higher is better
Time to first token
358 ms
lower is better
Efficiency
0.31 tok/s/W
Electricity cost
$0.151/Mtok
Power draw
331 W
AMD EPYC 7352 24-Core Processor (znver2) · 126 GB RAM · Linux · llamafile
gpu4.pics/r/ls-result-1

All benchmarks on this rig

ModelQuantGenPromptTTFTPowerPerf/W$/MtokScore
Meta Llama 3.1 8B InstructQ4_K_M1033,552358 ms331 W0.31 tok/s/W$0.151/Mtok1,008

How it stacks up — Meta Llama 3.1 8B Instruct

NVIDIA RTX 3090
103
NVIDIA H100
120
NVIDIA A100 80GB
110
NVIDIA RTX 3090 Ti
110
NVIDIA RTX PRO 6000 Blackwell Workstation Edition
103
NVIDIA RTX A6000
90.5