r/LocalLLaMA • u/ColorsOfCosmos • 2h ago
Discussion Kaitchup posted Qwen3.8 27B Benchmarks for quants from Q4 to Q1
https://kaitchup.substack.com/p/qwen38-27b-gguf-benchmark-q4-to-q1Kaitchup just posted results of his benchmarks for Qwen3.8 27B for quants from different labs, Q4 to Q1, .
All the details are hidden behind the paywall, but high level result is visible and looks like for people with 16GB cards UD Q3_K_XL is a winner - it has accuracy of 100% and size is only 12.8GB.
24
u/spaceman_ 2h ago
Paywalled, sad.
15
u/ColorsOfCosmos 2h ago
I think that the most important piece of data is open - the graph shows size and accuracy for every gguf.
6
u/rossimo 1h ago
The ISTA-DASLab quants are good. I've been using the Q3, and it's working well with reverse-engineering tasks.
5
u/fgk55555 1h ago
I was going to say, I've been really liking the ISTA-DASLab IQ3 quants. They just uploaded a new MTP model set and I'm going to re-download, but they're by far the best performing for the size. I can fit loads of context in my 16GB card, but even 12GB users could probably get the 27B now. They've done a great job.
1
u/rossimo 1h ago
I honestly use it on my 24gb vram card for the speed and the additional context room.
1
u/fgk55555 1h ago
Nice. I was really woe-is-me when the 27B dropped that I didn't have more VRAM for a better quant but after using the IQ3 a bunch I'm pretty satisfied. 120k context and 55tg on my $700 card is good enough. We can ride out the AI bubble now.
1
4
u/giri24343 2h ago
How are the ridge models ? Looks like the ridge is sitting at 12GB and still as good as q4.
2
u/stoppableDissolution 1h ago
Most definitely bs. Theres not a single chance a model trained in 16 bit loses literally nothing even at q8, let alone q3. It just means ulrasaturated benchmark.
1
u/Charming_Clothes8990 1h ago
has anyone tried that q3 quant for roleplay? wonder if the personality stays consistent or if it starts drifting.
1
1
u/ea_man 56m ago
Beware that it don't mean: how good is the model at coding, how good is at not losing details in a 130k coding session.
I means how close the benchmaxed result are similar to the benchmaxed result of f16 at zero context, which is a terrible way to quant a model.
There's a practical way to valuate those models: first have a SOTA or your best 27B analyze the tensor weighting and valuate how that should affect coding, ability to solve _new_ problems, follow specifications, persistence of code.
Then you test that: have a SOTA design a prompt to engage such properties and put those quants to the test, evaluate the coding session by a SOTA not one of the model that made those ofc, so you see how good are for such purpose.
1
u/ea_man 56m ago edited 49m ago
Beware that it don't mean: how good is the model at coding, how good is at not losing details in a 130k coding session.
I means how close the benchmaxed result are to the benchmaxed result of f16 at zero context, which is a terrible way to quant a coding model.
There's a practical way to valuate those models: first have a SOTA or your best 27B analyze the tensor weighting and evaluate how that should affect coding, ability to solve _new_ problems, follow specifications, persistence of details at long horizons.
Then you test that: have a SOTA design a prompt to engage such properties and put those quants to the test, evaluate the coding session (not fucking pelicans SVG!) by a SOTA not any of the models that made those ofc, (no freebee Gemini don't count either as it's dumb as rocks) so you see how good those are for real job.
1
u/derspenti 28m ago
Even taking the 100% with a shovel of salt, 12.8GB for a 27B means a 16GB card gets the model and real context, not one or the other. That alone makes Q3 worth trying.
2
u/pl201 2h ago
If the report stated 100% accuracy at q3, it is worthless to read.
7
u/James-Keydara 1h ago
It might be a good indicator of how benchmaxxed the model is because he used prompts from known benchmarks. It's still impressive that 100% recall is even possible at q3
1
u/Embarrassed_Soup_279 1h ago
101% to bf16... idk if i trust it. from my experience you definitely notice a difference between even Q4 K XL and Q5 K XL.
25
u/SnooPaintings8639 2h ago
If this would hold, then for Qwen 27B, Q3 is the new Q4.