r/LocalLLaMA 1h ago

Discussion CMP170Hx “Spark” Machine

I got the CMP170 cards and unlocked them. I wanted to share my set up for CUDA since maybe it would be useful to others.

First off, I hate e-waste and we are in a special time for RAM. I wanted to have a DIY CUDA box, and I had started by adding additional cards to an old asus predator prebuilt I had around, which also had 64gb DDR5. To add the CMPs I needed more CPU lanes and newegg had some really good deals on CPU/MB/etc combos. Didn’t need a combo with RAM, otherwise I would have gotten it in newegg microcenter.

Anyway, I got a cheap case, some noctua fans for the cards, and transferred the memory/ssds. Placed previously owned cards on oculink slots, and used the main x16 for the GPU switch that houses the two CMP170s, so their effective speed is 2x16 across and with the other cards (which are 4x4, and therefore same speed).

Qwen Flash Next, turns out, fits very nicely in these cards. There is also a repository for deepseek, but you’d need at least 3 64GB cards to run it, and with prices rising, it will be hard to justify the gamble of buying ex mining cards for LLMs.

However…so far, these cards are great. Concurrency is good, prompt processing averages 4000 tps on Flash Next, decode is 80+ on a single stream. No MTP added. Third picture shows the 3 models I am now running in this CUDA box (flash next, qwen 27b, gemma 26b).

Anyone else trying out Flash Next on these cards?

14 Upvotes

34 comments sorted by

4

u/Automatic_Two4291 1h ago

These cards seem awesome. If only they would stop rising in price

4

u/FullstackSensei llama.cpp 1h ago

Grab what you can, while you can. Prices of everything are going up again.

3

u/FullstackSensei llama.cpp 1h ago

Really cool setup!

Do you have some more pics of the GPU setup, especially the PLX switch board? Which one did you get and how do you find it? How are temps using these noctua fans? Are you power limiting the cards? Did you try a 120 or 140mm 3k rpm fan, or better yet something like the Arctic S12038?

I've been using the S8038-7k to cool my Mi50s and it's been really good. Has very high static pressure and can keep the cards under 60C even at it's idle 2k rpm with MoE models.

Sorry for the barrage of questions. Been thinking of getting one of those PLX boards to move my P40s.

2

u/Miserable-Dare5090 1h ago

Yes, to all the questions:
1. I got the ADT Link 2 GPU switch because it’s cheap ish (170 on aliex) with the stand and the pcie card x2 MCIO ports. Cards do native 2x16 through the switch, no extra config needed. Peer 2 peer is enabled, and tensor parallel works.

I just found this on AliExpress: https://a.aliexpress.com/_mOq5cs3

The PEX88096 is a better switch though, and has like 80 pcie lanes / 4-5 slots. However it’s 4X the price. YMMV

  1. I didn’t just power limit them; I tuned them and tested the memory with a github project (search for 170tune). I think the Russians were aware of this mod long before we were…Tuning it will allow you to overclock them. at 180W, NDIV 72, humming nicely. You can go to 200W safely and without long term concerns too, but 180W buys you 99% of the juice.

  2. You need cooling. I had larger fans that were louder, but I wanted to reduce the engine noise. Noctua fans and a pwm fan switch with a knob, I can turn them down to whisper quiet or
    crank to whoosh. At whoosh levels, cards stay 35-42C when RT is 35C (hot room). Without fans, the temps rose super fast to 100C when I first tested them. I would recommend fans. The little noctua fans are a fancy add on, but they really are the quietest fans I have found.
    These work well for now. Temps even during the overnight soak and memory integrity test did not rise above 50-55. The Ryzen processor, on an AIO cooler, rose way more!!

2

u/FullstackSensei llama.cpp 13m ago

180-200W is a lot lower than I expected for these cards. No wonder you don't need a lot of air to cool them.

The name of the game with these passive cards is static pressure. A fan that has high static pressure can push a lot more air through the heat sink even at low rpm. Conversely, low static pressure fans can scream all they want, but won't push much air through. That's why I mentioned the S8038, which I use, and the S12038. Even at idle they can push a lot more air through than any desktop fan at max rpm, noctuas included. Arctic also does a great job keeping noise low, not far behind noctua.

1

u/Miserable-Dare5090 3m ago

The ones I got from China sounded like a private airport, so they had to go. I do have a shroud for a 170mm and a 120mm fan I 3D printed. May look at the arctic fans to see if it’s worth it.

By the way. I have read your posts here for a year so I know you have a strix halo as well.

These fuckers are an amazing sidecar to the strix!! I was tempted to get a 3rd one for that…although, the R9700 seems like the best bet for that purpose right now.

2

u/Miserable-Dare5090 34m ago

The Switch is in the aliexpress site, I bought the one with a redriver, to allow full 16 lanes. Is it wasteful to use a pcie5 x16 slot? Probably. But the cards communicate stably, and the other cards are bearing their models well. I just got nvme to gen5x4 adapters to see if I can boost the oculinked cards’ interconnect more…and then I have to stop. I can’t keep buying shit for this stuff anymore or I will go broke!!

I think the Broadcom PEX 88096 switches are good, I would go for it if you have several cards you want to p2p on a single root pcie. It’s what the cmpunlock peeps recommend on discord, fwiw

1

u/FullstackSensei llama.cpp 18m ago

All my systems are now dual CPU, so have plenty of lanes, and having moved completely to data center cards, I now have p2p with stock drivers everywhere.

The thing with the P40s is, they're running with dual Broadwell, which while they're quite good when offloading layers to RAM, they don't have enough memory bandwidth to handle larger models. Skylake/cascade lake boards with lots of slots are still very expensive.

I've looked at both the PEX88096 and the PEX8796 (previous version, Gen 3). Both have 96 lanes, as the name suggests, but the latter is for some reason more expensive. I can find the 88096 boards for half the aliexpress prices, but it's 5 cards "only" and I have 8 P40s. There used to be an 8796 board with 7 slots, mix of x8 and x16, but it's become very hard to find.

Broadcom actually ruined the whole PCIe switch party when they acquired PLX, the startup that originally designed and made these switches. Before the acquisition, they were much cheaper and you could find them in so many enthusiast and workstation boards, but I digress.

2

u/cibernox 1h ago

I have two of those in the mail, eager to test them. I got a couple Arctic P12 pro PWM to cool them.

I intended to run this same qwen-flash, but for 3000$ this is probable the budget king setup for running models below 200B.

2

u/Miserable-Dare5090 1h ago

You can enable peer to peer, so get a switch to have the cards tensor parallel to each pther and not througj the pcie bus

1

u/cibernox 1h ago

But the most I can get is pcie 2.0 16x, isn’t that too slow for tensor parallelism? I kind of had made peace with the fact that I could only use pipeline parallelism.

2

u/WeAreSven 58m ago

That's basically gen 4 x4 speeds which lots of people have done TP with. gen 3x4 is too slow however, I know because my mobo is inconvenient in all of the wrong ways and I've been looking for workarounds.

1

u/Miserable-Dare5090 1h ago edited 56m ago

nope, running that flash next on TP2 right now. How it got enabled, you’ll have to ask Deepseek V4 Flash, my hard little worker. 13Gb/s bidi is enough, and skipping CPU/Host means the kind of latency you want for TP.

Now, unmodded 2x4? pipeline will work really well

2

u/cibernox 56m ago

Mine are modded already for 16x. In 10 days or so I’ll give it a go. Then I’ll check how pricey PCIe switches are.

1

u/Miserable-Dare5090 47m ago edited 39m ago

https://a.aliexpress.com/_mOgZ7cf 120, put the card passthrough pcie card on an x16 slot (the ryzen 9900x has a 2 core igpu so that allows me to use the gpu slot for this instead).

1

u/sooki10 1h ago

Really?

2

u/WeAreSven 57m ago

If that's 3k for 2 cards can you tell me where you got them? I'm looking to do exactly the same thing and the main chinese guy on ebay keeps raising his prices daily.

2

u/cibernox 54m ago

I got them in Alibaba probably by the Chinese guy you mention, but I got them the day before they went to 1899 and now 2099 but the seller honored the quote.

1

u/Miserable-Dare5090 44m ago

Alibaba, its a nail biter for us Americans who are used to being complainy little customers. But the cards came. Took like 15 days but they arrived, and were good. AILFond is the seller I used. I didn’t bother with the “bitcoin only, through whatsapp” vietnamese sellers. Not sure if its legit but that’s way too sketchy for me.

2

u/Kahvana 1h ago

Super cool setup! I do worry about the aging of these cards, assuming they have been heavily used.

3

u/sooki10 54m ago

The card is based on Nvidia  enterprise grade GA100 architecture, designed for data centre life.  Even the consumer grade 3090s of that gen still holdimg strong..  If you can get one from a place that has some degree of buyer protection and immediately test it hard, it isnt all that risky.

1

u/Miserable-Dare5090 59m ago

I agree, but more than the 3-4th string 3090 out there for same price? I saw someone cleaning vape fluid from the PCB in this sub…I mean, it’s slim pickings out there.

64G, 1.4Tb bandwidth, CUDA, Ampere/Sm80…It’s not getting cheaper. But I’ll post when they break :)

2

u/leonbollerup 38m ago

care to share your config.. i have those cards aswell...

1

u/Miserable-Dare5090 29m ago

Sure, what config? runtime config for flash next? TP2 or PP2? Ie are you using them
on full x16 lanes or nay

1

u/Miserable-Dare5090 18m ago

I asked my agent who was really the one who tweaked the config:
Reddit is not allowing me to post pictures, I DMed tou

1

u/Miserable-Dare5090 12m ago

You get 4 full context sessions from 2 cards:

```
vllm serve /models/Qwen3.8-Flash-Next-
AWQ-INT4
--served-model-name qwen38-flash-next-awq
-host 0.0.0.0 \
--port 8000 \
--load-format safetensors
--max-model-len 262144 \
--max-num-seqs 8 \
--gpu-memory-utilization 0.95 \
--tensor-parallel-size 2
--enable-prefix-caching
--enable-chunked-prefill
--max-num-batched-tokens 8192 \
-CC.cudagraph_mode-PIECEWISE\
-cc.splitting_ops-"[\"vllm:: unified_at
tention_with_output\i,
Ạvllm: :unified_mla_attention_with_out
put\", \ 'vllm: :mamba_mixer2\",
\'vllm: : mamba_mixer\',
\"vllm:: short_conv\",
\"vllm:: :qwen3_8_flash_next_ple_short_c
\"vllm: :qwen3_8_flash_next_qsa_with_ou
tput\", \'vllm: :linear_attention\", \"vllm:: qwen_gdn_attention_core\",
\"vllm:: qwen_gdn_attention_core_fused
norm_packed\",
\"vllm:: sparse_attn_indexer\"
\'vllm: :ple_mmap_lookup\"]" \
--no-enable-flashinfer-autotune
--disable-custom-all-reduce \
--kv-cache-dtype auto\
--enable-auto-tool-choice \
--tool-call-parser qwen3_xml \
--reasoning-parser qwen3 \

```

All credit goes to Hermes-on-Deepseek, not me. Homie (my bot) wanted you to know: “The -cc.* flags are the custom CUDA-graph config targeting the Flash-Next fused ops (unified attention, MLA, mamba/linear-attention, sparse-attn indexer, PLE lookup) with cudagraph_mode=PIECEWISE — that's the per-op graph split, and the --no-enable-flashinfer-autotune + --disable-custom-all-reduce pin single-host NVLink/PCIe behavior instead.”

1

u/quantgorithm 35m ago

the CMP cards are NOT 2x16 unless you physically modified them. You can only software unlock to, I believe, 2x4.

1

u/Miserable-Dare5090 21m ago

you can buy them modded. I did, so did several people who commented in this thread.

It needs additional capacitors soldered in, and it was part of the 1200 price I got them for.

So far, people can get them to identify as gen3 but no one has been able to show gen3 speeds. If it does happen, it will be sweet but they are already doing TP at 2x16.

1

u/Khipu28 1h ago

Are you sure those Noctua fans do anything? I think they don't have enough static pressure to push enough air through.

1

u/Miserable-Dare5090 1h ago edited 1h ago

Yeah, they keep the cards below 50C while running at full, 34C at baseline. No fans and cards will rise to 100C in 5 minutes.

And they don’t sound like an engine

1

u/beryugyo619 35m ago

OP has four of the 4cm ones, I guess they add up. Power draw certainly does

1

u/Miserable-Dare5090 1m ago

they’re running on a single sata cable man. both cards plus fans won’t hit 500W.

Compare to 3090s needed for 128gb (5? at 300 watts each, downvolted to 250…mmm maybe about 1250W?)