gpu4.pics
data via LocalScore
Apple M4 Max
llamafile 0.9.1
128 GB unified Apple
running Llama 3.2 1B Instruct· Q4_K_M· 1.5B
184
tokens / sec
1,336 · Excellent
Generation
184tok/s
higher is better
Prompt / prefill
3,893tok/s
higher is better
Time to first token
301 ms
lower is better
Efficiency
5.2 tok/s/W
Electricity cost
$0.009/Mtok
Power draw
36 W
Apple M4 Max 12P+4E · 128 GB RAM · Darwin · llamafile
gpu4.pics/r/ls-result-6

All benchmarks on this rig

ModelQuantGenPromptTTFTPowerPerf/W$/MtokScore
Llama 3.2 1B Instruct1.5BQ4_K_M1843,893301 ms36 W5.2 tok/s/W$0.009/Mtok1,336

How it stacks up — Llama 3.2 1B Instruct

Apple M4 Max
184
NVIDIA L40S
355
NVIDIA RTX 3090 Ti
354
NVIDIA H100
335
NVIDIA RTX 3090
330
NVIDIA RTX A6000
315
NVIDIA A100 80GB
308