gpu4.pics
data via LocalScore
Apple M4 Max
llamafile 0.9.1
128 GB unified Apple
running Meta Llama 3.1 8B Instruct· Q4_K_M
51.6
tokens / sec
250 · Passable
Generation
51.6tok/s
higher is better
Prompt / prefill
607tok/s
higher is better
Time to first token
1.99 s
lower is better
Efficiency
1.78 tok/s/W
Electricity cost
$0.029/Mtok
Power draw
32 W
Apple M4 Max 12P+4E · 128 GB RAM · Darwin · llamafile
gpu4.pics/r/ls-result-11

All benchmarks on this rig

ModelQuantGenPromptTTFTPowerPerf/W$/MtokScore
Meta Llama 3.1 8B InstructQ4_K_M51.66071.99 s32 W1.78 tok/s/W$0.029/Mtok250

How it stacks up — Meta Llama 3.1 8B Instruct

Apple M4 Max
51.6
NVIDIA H100
120
NVIDIA A100 80GB
110
NVIDIA RTX 3090 Ti
110
NVIDIA RTX PRO 6000 Blackwell Workstation Edition
103
NVIDIA RTX 3090
103
NVIDIA RTX A6000
90.5