r/LocalLLaMA 18h ago

Resources Which current local models that can run within 128GB generate the best SVG pelicans?

Post image

I used a famous Simon Willison's pelican riding a bicycle prompt on the biggest local LLMs that can run on 128GB Apple Silicon. U used quantizations by Unsloth.

Qwen3.8 Flash-Next gives a lot of details. DeepSeek V4 Flash is strangely underwhelming. Qwen3.8 27B still rocks, and I like its consistent minimalism.

Is Qwen3.8 27B still large at 31GB? It is! But for this tasks 2-bit quantizations (at around 12GB) will give the same results. For more complicated coding, 4-bit are more than enough. RTX cards are well enough!

See:

39 Upvotes

64 comments sorted by

86

u/jacek2023 llama.cpp 18h ago

In my opinion, that test doesn't make sense because the models were trained on that specific task. You should be more creative and try something different to avoid benchmaxxing

46

u/ManIkWeet 17h ago

Clear indicator is that every single pelican faces the same direction.

10

u/jacek2023 llama.cpp 17h ago

very good point

10

u/rditorx 14h ago

Left to right preference is common among people using the Latin alphabet as their primary alphabet. I guess it would also be common without benchmaxxing a pelican

11

u/enemyofaverage7 15h ago

tbh the direction thing is probably more likely because bicycle photos are typically taken from that side (generally referred as 'drive side') to display the parts that are on it.

1

u/ManIkWeet 15h ago

Most it seems, but not all. So I would expect most, but not all pelican SVGs the same :)

6

u/BigYoSpeck 9h ago

2

u/ManIkWeet 9h ago

I like your style, make up your own rules!

11

u/BigYoSpeck 16h ago

How about a pelican reading a bash script from a teleprompter?

6

u/jacek2023 llama.cpp 15h ago

Change pelican to llama

5

u/No_Advance3911 14h ago

7

u/jacek2023 llama.cpp 14h ago

see? now it's heading left ;) so it is not benchmaxxed

3

u/No_Advance3911 14h ago

Visually, Qwen3.8-27B-UD-Q2_K_XL is already a huge achievement :D

3

u/BigYoSpeck 13h ago

1

u/jacek2023 llama.cpp 12h ago

good local llama :)

1

u/Ok-Direction-4480 11h ago

Woah that looks beautiful

1

u/No_Advance3911 13h ago

He's just missing a blue suit and some yellow hair, then the resemblance would be way more obvious :D

5

u/quiteconfused1 17h ago

It's not the fact that content has been trained on ... It's the fact that the quant has degraded the feature or not.

That's what this test demonstrates - Quant advantage.

2

u/pmigdal 17h ago

Yeah, I know the phenomenon of pelicanmaxxing.

While usually pelicans are indicative of performance on other SVGs, I will try something more creative the next time.

1

u/kiwibonga 13h ago

Those are very disappointing results for "benchmaxxing" - look at those underwhelming pouches.

24

u/armeg 17h ago

1999: AI will cure cancer.

2026: AI will make pictures of pelican riding a bike.

10

u/hejj 16h ago

Using $10,000 worth of computer hardware

8

u/BlobbyMcBlobber 15h ago

10K?

Is this 2025 again?

6

u/SpicyWangz 15h ago

What if the real cancer was just the ai we made along the way

3

u/armeg 15h ago

lmao

8

u/OwnGear3892 18h ago

Thanks for sharing. I guess Deepseek V4 Flash under-performing is reasonable as it's run on IQ3, quite natural performance drop as trade off to fit in 128 GB unified ram.

1

u/pmigdal 14h ago

I seems that quantization hurts.

-1

u/uti24 17h ago

I guess Deepseek V4 Flash under-performing is reasonable as it's run on IQ3

Usually, a bigger model at a smaller quant should work better than a smaller model at a bigger quant, since the sizes here are comparable, and DeepSeek is even bigger, so the comparison is fair.

Still, the quants could be of different quality, and the Deepseek V4 could simply have less training data like that and more data for something else.

4

u/quiteconfused1 17h ago

This is speculative. At some point the quant will be degraded so much it will fail. Equally some models may be so compressed that quantizing will have exponential loss in contrast to smaller models.

Tldr there isn't 1 rule, and it's why the pelican riding a bike is important.

2

u/uti24 16h ago

This is speculative. At some point the quant will be degraded so much it will fail.

Sure, but usually that will happen after the size of the bigger quantized model becomes smaller than the size of the smaller, less-quantized model (and often even then bigger model stays better). Here, the total size of DeepSeek Flash is 104 GB, while Qwen Flash is 94 GB.

Also, bigger models are much more resilient to quantization. They may lose some precision, but their reasoning tends to suffer less.

8

u/jaegernut 17h ago

Can we try a different animal next time

5

u/gh0stwriter1234 12h ago edited 12h ago

A rhino on a dino, Qwen 3.8 Flash Next Q4_XS unsloth on 2x MI50 low reasoning

1

u/Ok-Direction-4480 10h ago

Looks a bit janky but still nice.

1

u/gh0stwriter1234 9h ago

Definitely janky I think some of that is to blame on low reasoning.

10

u/bonobomaster 18h ago

4

u/pmigdal 16h ago

Wow, this is awesome! Will give it a try

2

u/SpicyWangz 15h ago

I’m already fully weevilmaxxed

6

u/mickabrig7 17h ago

A visual LLM test showing several retries including failed attempts ? Am I in heaven ?

5

u/ixdx 16h ago

Qwen3.8-27B-MTP-Q6_K 21.8 GiB with bartowski imatrix --reasoning-effort xhigh

I made several attempts. The second one even turned out to be animated (rotating wheels and simulated airflow).

2

u/pmigdal 15h ago edited 14h ago

Q8 is close to lossless, regardless of provider. I used medium effort, and I guess it is behind the difference.

For `xhigh` effort it gets similarish.

1

u/lhg31 14h ago

Are you sure you used medium for flash next? The amount of tokens suggests it was xhigh.

2

u/pmigdal 14h ago

I mean, I used medium Qwen3.8 27B.

4

u/insu_na 17h ago

Wait, your Qwen3.8 Flash-Next stops thinking at some point? :O

5

u/ArrogantAnalyst 17h ago

Thank you! I was looking for a pelican optimized model.

3

u/2muchnet42day Llama 3 16h ago

Exactly. Thank you very much

2

u/geneusutwerk 17h ago

Definitely top right

2

u/Sad_Recording_1290 17h ago

Are the pelicans really the benchmark now?

2

u/AleksandrNikitin 16h ago

more pelicans for the pelican god

2

u/mailto_devnull 14h ago

So 3.8 Flash-Next burns through even more tokens than 3.8 27B.

That's not a good trend.

1

u/Hannibalj2ca 17h ago

is that what people make with large language models, cartoon pelicans? I say, yes!

1

u/my_name_isnt_clever 12h ago

I had Qwen 3.8 Flash Next vibe up a little reiterative SVG making script where it generates a SVG then self-corrects any minor flaws before the final output, I've been pretty impressed with what it can do with novel prompts.

1

u/Squidgical 11h ago

The best test for an LLM is a test no one has ever heard of.

The worst test is one everyone has heard of, and that the LLM definitely has specific training for.

1

u/Ok-Direction-4480 11h ago

What do you mean Qwen 3.8 27B used the fewest tokens?

1

u/_supert_ 11h ago

Now show me a bike riding a pelican.

1

u/Ok-Direction-4480 10h ago

Just a question, why SVG? Aren't image generators (especially fine-tuned ones for cartoons) more efficient?

1

u/BigYoSpeck 9h ago

People are quick to jump to the conclusion of the training data being contaminated by this "test". First, I don't imagine SVG creating ability is remotely a focus of the training, especially not specifically the pelican

Secondly, that doesn't account for the ability to structure SVG for things they will never have been asked to do before:

I know this is a little disjointed compared to the almost pixel perfect SVG they can create when it's an ambiguous prompt they are free to make assumptions on. But being able to oneshot from a photo and largely maintaining the positioning and vibe is still insane

I honestly believe their SVG creation capability is fundamentally just a byproduct of their raw coding ability. Being able to create styled UI components in general demands skill with positioning and composition. So I don't think they are benchmaxed on SVG, they are just very capable in the domain that lends well to making SVG

1

u/No_Dragonfruit_8651 8h ago

Who gives a hairy rats cock about pelicans Im not sure

1

u/simrankoulsm 4h ago

Nice comparison. I would be curious to see SVG validity, render success, token count, latency, and editability measured alongside visual quality. For local use, the best model may not be the one with the prettiest one-shot pelican, but the one that reliably emits valid, compact SVG that survives small prompt edits and quantization.

1

u/vulcan4d 1h ago

The next model will be trained on pelicans and bikes

0

u/Fun_Jaguar8231 17h ago

Oh f**k me with those pelicans