gpu4.pics
data via LocalScore
NVIDIA GeForce RTX 4060 Ti
llamafile 0.9.2
16 GB VRAM NVIDIA
running Gemma 3 4b It· Q4_K_M· 4.6B
74.4
tokens / sec
1,113 · Excellent
Generation
74.4tok/s
higher is better
Prompt / prefill
4,885tok/s
higher is better
Time to first token
264 ms
lower is better
Efficiency
0.63 tok/s/W
Electricity cost
$0.075/Mtok
Power draw
118 W
AMD EPYC 7352 24-Core Processor (znver2) · 126 GB RAM · Linux · llamafile
gpu4.pics/r/ls-result-14

All benchmarks on this rig

ModelQuantGenPromptTTFTPowerPerf/W$/MtokScore
Gemma 3 4b It4.6BQ4_K_M74.44,885264 ms118 W0.63 tok/s/W$0.075/Mtok1,113