My 160GB local LLM rig (4x V100 + 4x 3090)
gpu4.pics
via TrifleHopeful5418
NVIDIA Tesla V100
4× NVIDIA Tesla V100 16GB + 4× NVIDIA RTX 3090 · 160 GB total VRAM · AMD Threadripper
32 GB VRAM NVIDIA 300 W TDP
running Qwen3 235B· Q4
15
tokens / sec
Generation
15tok/s
higher is better
Prompt / prefill
—tok/s
higher is better
Time to first token
—
lower is better
Composite score
—
Model size
235B
Runtime
—
All benchmarks on this rig
| Model | Quant | Gen | Prompt | TTFT | Power | Perf/W | $/Mtok | Score |
|---|---|---|---|---|---|---|---|---|
| Qwen3 235B | Q4 | 15 | — | — | — | — | — | — |