gpu4.pics
data via LocalScore
NVIDIA RTX 6000 Ada Generation
48 GB VRAM NVIDIA
running Meta Llama 3.1 8B Instruct· Q4_K_M
51.3
tokens / sec
1,038 · Excellent #12 of 370
Generation
51.3tok/s
higher is better
Prompt / prefill
5,487tok/s
higher is better
Time to first token
252 ms
lower is better
Composite score
1,038
Model size
8B
Runtime
gpu4.pics/r/ls-nvidia-rtx-6000-ada-generation-meta-llama-31-8b-instruct-q4-k-m

All benchmarks on this rig

ModelQuantGenPromptTTFTPowerPerf/W$/MtokScore
Meta Llama 3.1 8B InstructQ4_K_M51.35,487252 ms1,038

How it stacks up — Meta Llama 3.1 8B Instruct

NVIDIA RTX 6000 Ada Generation
51.3
NVIDIA H100
120
NVIDIA A100 80GB
110
NVIDIA RTX 3090 Ti
110
NVIDIA RTX PRO 6000 Blackwell Workstation Edition
103
NVIDIA RTX 3090
103
NVIDIA RTX A6000
90.5