Modder Adds 32GB Tesla V100 to Gaming PC, Running 27B AI Model at 32 Tokens/Second
Updated
Updated · Wccftech · Jul 26
Modder Adds 32GB Tesla V100 to Gaming PC, Running 27B AI Model at 32 Tokens/Second
1 articles · Updated · Wccftech · Jul 26
Summary
Tymscar fitted an NVIDIA Tesla V100 into an RTX 4080-based gaming PC, giving the system 32GB of usable VRAM for local AI workloads rather than gaming.
A £200 setup—about $266 for two 16GB V100s and an SXM2-to-PCIe adapter—solved the VRAM bottleneck that had limited larger language models on the existing GPU.
The modified system ran Qwen3.6-27B-MTP quantized at 19GB with a 128K-token context at 32 tokens per second, while prompt processing reached 133 to 160 tokens per second.
Getting there required custom cooling and power work because the Tesla V100 lacks a standard PCIe slot, display outputs and PCIe power connectors; its stock cooler also hit 82 dB before fan-speed tweaks.
The build highlights a low-cost path to offline home AI inference, using older data-center GPUs that can still deliver 900GB/s bandwidth for small to mid-sized models.