gpu4.pics
measured on this rig
Apple M3 Max
MacBook Pro — the gpu4.pics dev machine
128 GB unified Apple
running IBM Granite 4.1 8B· Q4_K_M
57.6
tokens / sec
593 · Strong
Generation
57.6tok/s
higher is better
Prompt / prefill
481tok/s
higher is better
Time to first token
133 ms
lower is better
Composite score
593
Model size
8B
Runtime
ollama
Apple M3 Max · 128 GB RAM · macOS · ollama
gpu4.pics/r/apple-m3-max-128-granite

All benchmarks on this rig

ModelQuantGenPromptTTFTPowerPerf/W$/MtokScore
IBM Granite 4.1 8BQ4_K_M57.6481133 ms593