M3 Ultra 512GB does 18 t/s with DeepSeek R1 671B Q4

gpu4.pics
via AliNT77
Apple M3 Ultra
1× Apple M3 Ultra (512GB unified memory) · 512 GB total VRAM · Apple M3 Ultra
512 GB unified Apple
running DeepSeek R1 671B· Q4
18
tokens / sec
Generation
18tok/s
higher is better
Prompt / prefill
tok/s
higher is better
Time to first token
lower is better
Composite score
Model size
671B
Runtime
Apple M3 Ultra · 512 GB RAM
gpu4.pics/r/curated-1j8r2nr

All benchmarks on this rig

ModelQuantGenPromptTTFTPowerPerf/W$/MtokScore
DeepSeek R1 671BQ418

How it stacks up — DeepSeek R1 671B

Apple M3 Ultra
18
NVIDIA RTX 3080
7.5