r/LocalLLaMA 15h ago

Discussion Deceptive model quantization from AtomicChat?

I kept seeing guys in this sub saying how AtomicChat's Qwen3.8-Flash-Next quant is so good, fits in their machine when unsloth's can't, runs faster than other quants etc, so I went check out what's happening there.

First thing I noticed was that AtomicChat's Q4_K_M quant is suspiciously small when the ngram table is removed (only ~56GB), it seems like most of the tensors in this quant are IQ2_S instead of the usual Q4_K, Q5_K and Q6_K that you usually find in Q4_K_M quants, the GGUF filetype metadata also says IQ2_S instead of Q4_K_M. In their model card, their Q4_K_M also has suspiciously high KLD (0.084).

It seems pretty obvious to me that they're pretending a IQ2_S quant as a Q4_K_M, but at the same time I'm genuinely not sure because it can't be only me who found this right? How can nobody be pointing this out? Am I missing something or what may they be doing?

Their HF repo ID: AtomicChat/Qwen3.8-Flash-Next-GGUF

61 Upvotes

38 comments sorted by

View all comments

28

u/lhg31 15h ago

Well, they DO explain this, don't they?

Naming

Files are named by their measured bits per weight. A build whose expert tensors are IQ1_M is not a 1-bit model when the n-gram table sits at 6 bits and ffn_down_exps at 4.5; the real average is 3.84. The canonical type in the filename is the closest standard type by that average, so tooling can still detect it. For AD-4.27bpw:

Group Type Share of file Contribution
n-gram table Q5_1 41% 1.74 bpw
ffn_gate/up_exps IQ2_S, IQ3_S at the band 29% 1.24 bpw
ffn_down_exps IQ4_NL 24% 1.03 bpw
everything else Q8_0 5% 0.23 bpw

15

u/po_stulate 15h ago

I mean, so the quants that keep the ngram table unquantized and everything else IQ2_S should probably be called Q8_0 instead since the average bpw is closer to that? Makes no sense to me.

6

u/lhg31 14h ago

Mixed precision / Dynamic quants were always like that.

If you take gemma-4-26B-A4B-it-UD-IQ4_XS from Unsloth, almost 50% of their tensors are IQ3_S, but they also have IQ4_NL and Q8_0, so they average to IQ4_XS.

AtomicChat decide to include n-gram table in the equation. You may not like this decision, just like some people may not like Unsloth calling a model with half IQ3_S tensors a IQ4_XS. But they are open about it being mixed/dynamic, that's what matters.

Since n-gram table can sit on ssd, I don't care much about its quantization. So FOR ME, n-gram should be left out of the equation. But it's their quant and they were honest on their formula, so I'm good with it.

0

u/po_stulate 13h ago

That makes sense, but in my own opinion seeing a lot of IQ3_S for a IQ4_XS quant is actually exactly as expected, but half expert weights quantized to IQ2_S for a Q4_K_M not so much.