gpu4.pics
data via LocalScore
Apple M4 Max
llamafile 0.9.1
128 GB unified Apple
running Llama 3.2 1B Instruct· Q4_K_M· 1.5B
156
tokens / sec
378 · Passable
Generation
156tok/s
higher is better
Prompt / prefill
680tok/s
higher is better
Time to first token
1.97 s
lower is better
Efficiency
3.64 tok/s/W
Electricity cost
$0.014/Mtok
Power draw
46 W
Apple M4 Max 12P+4E · 128 GB RAM · Darwin · llamafile
gpu4.pics/r/ls-result-8

All benchmarks on this rig

ModelQuantGenPromptTTFTPowerPerf/W$/MtokScore
Llama 3.2 1B Instruct1.5BQ4_K_M1566801.97 s46 W3.64 tok/s/W$0.014/Mtok378

How it stacks up — Llama 3.2 1B Instruct

Apple M4 Max
156
NVIDIA L40S
355
NVIDIA RTX 3090 Ti
354
NVIDIA H100
335
NVIDIA RTX 3090
330
NVIDIA RTX A6000
315
NVIDIA A100 80GB
308