r/LocalLLaMA 2h ago

Discussion Kaitchup posted Qwen3.8 27B Benchmarks for quants from Q4 to Q1

https://kaitchup.substack.com/p/qwen38-27b-gguf-benchmark-q4-to-q1

Kaitchup just posted results of his benchmarks for Qwen3.8 27B for quants from different labs, Q4 to Q1, .

All the details are hidden behind the paywall, but high level result is visible and looks like for people with 16GB cards UD Q3_K_XL is a winner - it has accuracy of 100% and size is only 12.8GB.

59 Upvotes

22 comments sorted by

25

u/SnooPaintings8639 2h ago

If this would hold, then for Qwen 27B, Q3 is the new Q4.

8

u/Bulky-Priority6824 2h ago

Makes sense. I went from 3.6 27b q8 to 3.8 27b q6 and it's a big step forward with what I'm able to produce with it 

13

u/unia_7 1h ago

Well, Q6 is known to have almost no degradation relative to Q8. Even Q5 has very little degradation, as far as I know.

The question has always been about Q4 and Q3.

1

u/Bulky-Priority6824 1h ago

I don't know maybe but I could def tell a difference no benchmarks all vibes tho. Like driving a car except cars don't have logs that tell you they fucked up lol

24

u/spaceman_ 2h ago

Paywalled, sad.

15

u/ColorsOfCosmos 2h ago

I think that the most important piece of data is open - the graph shows size and accuracy for every gguf.

6

u/rossimo 1h ago

The ISTA-DASLab quants are good. I've been using the Q3, and it's working well with reverse-engineering tasks.

5

u/fgk55555 1h ago

I was going to say, I've been really liking the ISTA-DASLab IQ3 quants. They just uploaded a new MTP model set and I'm going to re-download, but they're by far the best performing for the size. I can fit loads of context in my 16GB card, but even 12GB users could probably get the 27B now. They've done a great job.

1

u/rossimo 1h ago

I honestly use it on my 24gb vram card for the speed and the additional context room.

1

u/fgk55555 1h ago

Nice. I was really woe-is-me when the 27B dropped that I didn't have more VRAM for a better quant but after using the IQ3 a bunch I'm pretty satisfied. 120k context and 55tg on my $700 card is good enough. We can ride out the AI bubble now.

1

u/LetsGoBrandon4256 transformers 1h ago

Interesting. Thanks for sharing.

4

u/giri24343 2h ago

How are the ridge models ? Looks like the ridge is sitting at 12GB and still as good as q4.

2

u/stoppableDissolution 1h ago

Most definitely bs. Theres not a single chance a model trained in 16 bit loses literally nothing even at q8, let alone q3. It just means ulrasaturated benchmark.

1

u/Charming_Clothes8990 1h ago

has anyone tried that q3 quant for roleplay? wonder if the personality stays consistent or if it starts drifting.

1

u/Ok-Buffalo2450 1h ago

Anyone that can share the paywalled content? Maybe archived?

1

u/ea_man 56m ago

Beware that it don't mean: how good is the model at coding, how good is at not losing details in a 130k coding session.

I means how close the benchmaxed result are similar to the benchmaxed result of f16 at zero context, which is a terrible way to quant a model.

There's a practical way to valuate those models: first have a SOTA or your best 27B analyze the tensor weighting and valuate how that should affect coding, ability to solve _new_ problems, follow specifications, persistence of code.
Then you test that: have a SOTA design a prompt to engage such properties and put those quants to the test, evaluate the coding session by a SOTA not one of the model that made those ofc, so you see how good are for such purpose.

1

u/ea_man 56m ago edited 49m ago

Beware that it don't mean: how good is the model at coding, how good is at not losing details in a 130k coding session.

I means how close the benchmaxed result are to the benchmaxed result of f16 at zero context, which is a terrible way to quant a coding model.

There's a practical way to valuate those models: first have a SOTA or your best 27B analyze the tensor weighting and evaluate how that should affect coding, ability to solve _new_ problems, follow specifications, persistence of details at long horizons.
Then you test that: have a SOTA design a prompt to engage such properties and put those quants to the test, evaluate the coding session (not fucking pelicans SVG!) by a SOTA not any of the models that made those ofc, (no freebee Gemini don't count either as it's dumb as rocks) so you see how good those are for real job.

1

u/derspenti 28m ago

Even taking the 100% with a shovel of salt, 12.8GB for a 27B means a 16GB card gets the model and real context, not one or the other. That alone makes Q3 worth trying.

2

u/pl201 2h ago

If the report stated 100% accuracy at q3, it is worthless to read.

7

u/James-Keydara 1h ago

It might be a good indicator of how benchmaxxed the model is because he used prompts from known benchmarks. It's still impressive that 100% recall is even possible at q3

1

u/Embarrassed_Soup_279 1h ago

101% to bf16... idk if i trust it. from my experience you definitely notice a difference between even Q4 K XL and Q5 K XL.