r/LocalLLM • u/Affectionate-Bed3439 • 13h ago
Question Dual RTX 6000 Threadripper build
I am considering getting this build for my work. I have a budget of around 60k. Mostly will be running a mixture of small models like qwen 3.8 27b, with expectation to also be able to run oversized models like GLM 5.3 at quants like Q4.
Do you have any thoughts or recommendations different options?
Here is the parts list I am looking at:
AMD Ryzen Threadripper PRO 9985WX — 64C/128T — B&H — $7,894.00
ASUS Pro WS WRX90E-SAGE SE — B&H — $1,299.99
2× PNY NVIDIA RTX PRO 6000 Blackwell Max-Q — 96GB ECC GDDR7 — B&H — $16,999.99 each / $33,999.98 total
TEAMGROUP T-Create Master 384GB — 8×48GB DDR5-6000 ECC RDIMM — Newegg — $10,199.99
2× WD_BLACK SN850X 8TB PCIe 4.0 NVMe SSD — Best Buy — $1,699.00 each / $3,398.00 total
Samsung 990 PRO 2TB PCIe 4.0 NVMe SSD — Best Buy — $389.99
MSI MEG Ai1600T PCIE5 — 1600W 80+ Titanium PSU — B&H — $619.00
Sliger CX4200a 4U Rackmount Chassis — Sliger — $289.00
Asetek 836SA-M1 360mm Threadripper/TR5 AIO — Sliger — $250.00
3× Noctua NF-A12x25 PWM 120mm Fans — Sliger — $75.00 total
Sliger GDRAIL-20XX-B General Devices 20" Rack Rail Kit — AVADirect — $123.04
CyberPower PR1500LCD Smart App Sinewave UPS — 1500VA / 1500W — B&H — $700.95
Ubuntu 24.04 LTS — $0.00
Total: $59,238.94 before tax (tax isn't real, tax can't hurt us (shhhhhh let me live in delerium))
6
4
u/Maplesyrup000 11h ago
Decent build but at $60k you could do better than a couple Pro 6000s. They were already somewhat overpriced at $8k and now are half your build price, while a $100k DGX Station with a GB300 has 288GB of HBM3e and ~8TB/s of bandwidth, so even an out of the box station from Nvidia is much better value.
3
u/Maplesyrup000 11h ago
Imo you should take a look at AMD too. ROCm 10 is a big upgrade and you’re paying a huge tax for CUDA and AMD has closed the gap significantly there. Might be a squeeze but you could get probably get 2 x AMD MI 355 Ps. They are around $18k from this vendor but you can find cheaper. They have 144GB HBM3e each and 4TB/s of bandwidth. You could more than double your bandwidth and get 50% more VRAM for nearly the same price.
1
u/qqeyes 5h ago
Can you find a gb300 system for less than 150k delivered before next calendar year? I have found it impossible to find one at $100k
1
u/Maplesyrup000 5h ago
Here you go: https://smttr.com/products/nvidia-dgx-station-1x-nvidia-gb300-grace-blackwell-ultra-desktop-superchip
In stock at $94k
Another one: https://configurator.exxactcorp.com/configure/VWS-158270643
1
u/qqeyes 5h ago
At this price would it be better to build a 4x rtx 6000 pro system? It would give more headroom for larger models? Appears most of the memory on the gb300 dgx station is ddr5 which is going to tank decode speed if he exceeds the vram budget?
1
u/Maplesyrup000 5h ago edited 3h ago
The NV Link, power efficiency and memory bandwidth are desirable, but you are right there are better value for money options than an out of box solution.
https://youtu.be/qV_K0nTF6gY?is=gr0KSG5vEOztv2_z
And to your point about multiple Pro 6000s, that’s not a great value for money since MSRP double compared to what it was 6 months ago. You can buy actual AMD datacenter grade cards with HBM3 as I mentioned in my other comment and those would smoke the 6000s in terms of VRAM capacity and bandwidth.
5
u/Racer4711 LocalLLM 10h ago
why such an expensive ryzen? 64 cores don't help with anything. same for the ram. my 4x rtx 6000 system has only
16 cores and 64gb ram and runs perfectly.
6
u/Gromann7 13h ago
I’d bump your Ubuntu version to 26.04. It’s a solid build for someone with a ridiculous budget. What else are you using it for? That particular threadripper is overkill if this is just an inference server
6
u/shout4 12h ago

Almost the same build in specs:
AMD Ryzen Threadripper PRO 9985WX — 64C/128T 512 DDR5 ECC 5600
AMD Ryzen Threadripper PRO 9965WX - 24C/48T 512 DDR5 ECC 5600
2x ASUS Pro WS WRX90E-SAGE SE
4× PNY NVIDIA RTX PRO 6000 Blackwell Workstation — 96GB ECC GDDR7 (top)
4x Gigabyte 5080's (bottom)
I personally like the open frame format, better heat dissipation, and I can use the more powerful RTX Pro 6000 WS GPU's.
2
u/hammeredhorrorshow 12h ago
You need 240v for your power supply?
3
u/DustNearby2848 12h ago
If this is for a business then apply for nvidia inception. They give you 10% off 6000 pros
2
u/meldas 8h ago
Just to share my experience, I have a third RTX Pro 6k arriving this week to give me enough room to run some of the more recent models like GLM 5.3 Flash.
There was a specific period of time when I was getting LLM fomo when I was not able to run Minimax M3 or MiMo V2.5, and similar LLM fomo happening again with GLM 5.3 Flash. 2 GPU is only barely able to fit these models, or not at all, and even then you're trading off concurrency, KV cache size, MTP, etc to get it to fit within VRAM.
At the current market price, you should also consider 4x RTX Pro 5k 72GBs, which can be bought for around 9K USD each, slightly better value per VRAM, and not too far off of your current GPU budget. Lower memory bandwidth, but TP4 vs TP2 you may be able to win back some performance. Also might be a pretty tempting setup when serving smaller models like Qwen 27B, where you have one GPU serving the small model for the multimodal capabilities, and you still have 3 other GPUs that can run a larger model like DSV4F, quite a common setup that I see in the blackwell discord
1
u/OvertaxedOne 8h ago
How many concurrent streams and what models? 192GB puts you into the realm of running the big boys but if you're shooting for high concurrency, you could still be pretty constrained with this build. If you're only trying to serve at very high speed to a small number of users, this would be a pretty reasonable build. Qwen3.8 27B should scream on this rig and be able to support dozens of concurrent streams. The bigger models, you could have issues, 192GB is kind of an uncomfortable size, you can't run a great quant of a big model, too big for small models even with higher concurrency. If you get a 3rd Pro 6000 now you're into QwenNext range with good concurrency and blistering TPS.
As others said, no reason for so much RAM, I'd cut that in half. Also no reason at all for the CPU, get the cheapest Threadripper Pro you can find that unlocks all the lanes. You also don't need all those SSDs unless this will be a dual purpose system, a few TB should be more than enough or, if you're doing this at work, just hook up 10G and store additional models on your NAS/file server.
1
u/mxmumtuna 8h ago
A DGX Station at $80k with Inception (or $90k without) would be considerably better.
1
u/blackbird2150 6h ago
For what it’s worth, I just bought my second 6000.
Find a better price on your Rtx. Microcenter has them for $13,800 in many markets still like mine. It’s enough of a difference you could fly out to buy it. Also choose one in a sales tax free state.
That being said, you should know that this gives you access to much better models, but glm 5.3 flash at a reasonable quant is a very low context fit, like 100k tokens?
Deepseek (with new vision adapter) same thing but cache is more optimized so you get a bit more. Normal deepseek is better but not perfect for large or concurrent processes.
Qwen flash next. No problem. Fp8 lots of context.
You’re overbuying on ram and cpu. You just don’t need it imo. I have a 9900x and 64gb of ram. And the one time I wanted more ram for flash-next, it allowed me to offload to ssd for minimal performance hit. Not sure if you save enough but considering dropping that and getting a third 6000.
You also don’t need an 8tb ssd. Models are a few hundred gigs unless it’s being used for more the a models and repo (like media). I have a 2x2tb and i have as many models as I could want to run with 2+ tb free.
1
u/Puzzleheaded_Base302 2h ago
you don't need to spend so much to just run qwen3.8-27b. get a RTX PRO 6000 and a whatever motherboard, CPU combo with minimum system RAM like 16GB, you are good to go. don't go with PCIe3.0 era motherboard/cpu. There are compatibility issues prior to PCIe4.0. This will set you back by $20K.
you absolutely do not need a fancy CPU to run inference.
1
u/Affectionate-Bed3439 2h ago
If I was running one instance of qwen 3.8 27b, sure. But I’ll be running multiple in parallel. I also will be doing ram offloading for massive models like GLM5.3 so the CPU is part of the inference machine as well
1
u/JustTooKrul 2h ago
Interesting to see the comments--I would have thought that 192GB of VRAM is an odd breakpoint and that additional RAM would allow you to run the next group of models, albeit slowly. But, if this is a production machine then maybe it would simply be too slow to run any inference using RAM?
0
u/dupontping 10h ago
Why do I feel like this is bait or just reddit bs
If you have a high enough budget, just google or write a prompt for what you want and your needs plus budget.
Trying to flex and get validation on reddit is not it.
1
u/Affectionate-Bed3439 10h ago
Because I’ve already done those and am looking for more feedback before spending this much money???? Am I crazy for doing that???
-2
u/Lotrimous 11h ago
If you can afford 2x rtx pro 6000, switch to a motherboard that will support 6x native pciex5.0 x16 slots & go with 6x rtx 5090 (card price would be similar, vram bandwidth is far superior)...
-2
u/Big_Booty_Pics 9h ago
Other than the "cool" factor and that this is the local llm subreddit, is there specifically a need to host locally? $60k is years of Claude max subscriptions which would provide you better usability and flexibility IMO.
1
u/Maplesyrup000 5h ago
A lot of people want to use LLMs for development or analytics but have sensitive data that cannot leave on-prem. Not everyone trusts OpenAI or Anthropic with their HIPPA data or other PII
6
u/conifer_v11 13h ago
two 96gb 6000s already hold glm 5.3 q4 with room. the 384gb dimm kit is the bloated line.
drop to 256gb ecc and one 8tb nvme. spend the leftover on a third 6000 if you actually want tensor parallel later, not more host ram.