My 160GB local LLM rig (4x V100 + 4x 3090)

gpu4.pics
via TrifleHopeful5418
NVIDIA Tesla V100
4× NVIDIA Tesla V100 16GB + 4× NVIDIA RTX 3090 · 160 GB total VRAM · AMD Threadripper
32 GB VRAM NVIDIA 300 W TDP
running Qwen3 235B· Q4
15
tokens / sec
Generation
15tok/s
higher is better
Prompt / prefill
tok/s
higher is better
Time to first token
lower is better
Composite score
Model size
235B
Runtime
AMD Threadripper · 256 GB RAM
gpu4.pics/r/curated-1l5wxoa

All benchmarks on this rig

ModelQuantGenPromptTTFTPowerPerf/W$/MtokScore
Qwen3 235BQ415

How it stacks up — Qwen3 235B

NVIDIA Tesla V100
15
AMD Instinct MI50 32GB
58