r/LocalLLaMA • u/de4dee • 20d ago
New Model Qwen3.8-2.4T-A95B Released
https://huggingface.co/Qwen/Qwen3.8-2.4T-A95B166
u/lm-gtfy 20d ago
Finally. What is the meaning of life?
127
38
u/xaeru 20d ago
→ More replies (1)27
u/tripazardly 20d ago
Oh no, not again.
11
u/Caffdy 20d ago
OOTL, what is that?
17
u/corysama 20d ago
42 is famously a joke from Hitchhiker's Guide to the Galaxy. A bowl of petunias is less famously the same.
HHGttG if full of jokes that build up over long stretches (multiple books sometimes) and punchline in a single sentence. Can't recommend enough.
3
11
u/Paganator 20d ago
Many people have speculated that if we knew exactly why the bowl of petunias had thought that we would know a lot more about the nature of the universe than we do now.
→ More replies (4)6
u/feelspeaceman 20d ago
It's still good for distilling to train smaller models to be smarter, basically the popular 27B was distilled from this giant 2.4T.
A good mother model will likely give birth to more capable smaller models, but it's still requiring a tons of trials and errors, those that costs money to retry.
217
u/Intelligent_Ice_113 20d ago edited 20d ago
what is the knowledge cutoff date?
299
91
u/JumpingJack79 20d ago
I kinda like early cutoff dates. Whenever a model tells me its cutoff is 2024, I'm like "Dude, it's 2026 now and you won't believe what's happened..." 😏
129
u/-dysangel- 20d ago
"The user is talking about a hypothetical timeline"
21
u/JumpingJack79 20d ago
Lol, yes. And to be fair, I'd be skeptical too if I was in their place hearing those things.
20
19
u/Mickenfox 20d ago
In a few decades people are going to be doing 2020s roleplay with ancient models.
→ More replies (2)2
u/that_one_guy63 20d ago
Would be interesting to train on only text before 1970. Or like way earlier. It would be so confused.
3
u/nabagaca 19d ago
There are people doing that for 1800s london text https://www.reddit.com/r/LocalLLaMA/comments/1qaawts/llm_trained_from_scratch_on_1800s_london_texts/
6
→ More replies (1)44
u/teleprax 20d ago edited 20d ago
One of my favorite LLM memories was telling GPT (5 maybe?) a bunch of stuff like elon doing the nazi salute and some other stuff going on around that time, and then it proceeds to lecture me on fake news and how none of that is remotely correct and how it would be super serious.
Then I tell it to do a web search and it starts it's next reply with:
"Ok. wow."
EDIT: Ok i found it, its not quite the same as I remember, but close enough
https://i.imgur.com/Ta33JIT.jpeg https://i.imgur.com/kPDa9tm.jpeg
25
u/personahorrible 20d ago
"Sounds like a bizarre Black Mirror meets Idiocracy arc."
Bruh. You have no idea.
15
14
u/OttoRenner 20d ago
Mine went bonkers after I showed it several screenshots and links of GPU prices...was very funny to witness 😂😂
11
45
u/Ok_Ocelot2268 20d ago
Cutoff: April 2026. Startoff: June 5, 1989
14
u/thepaligator 20d ago
I remember June 5th. Not a lot of foot traffic that day. Clear Streets. It was nice.
10
8
u/goldrunout 20d ago
What is a start off date? Do they not train on older sources?
23
u/Anwar6969 20d ago
that’s a tiananmen square joke. nothing happened on june 4th 1989
21
u/FastDecode1 20d ago
[ Removed by the CCP ]
12
u/LawfulLeah 20d ago
the fact that the comment you replied to was removed by reddit makes this 10x funnier
6
u/Anwar6969 20d ago
crazy shit, i had to appeal to lift off the warning and the comment ban. i apparently got flagged for rule 1
→ More replies (1)97
u/FullstackSensei llama.cpp 20d ago
Because, that's the only thing holding you from running a 5TB model?
25
u/fullup72 20d ago
You can always quantize to 1 bit and stream experts from a 5400rpm HDD.
12
u/FullstackSensei llama.cpp 20d ago
Why settle for 5400rpm when you can get old drives that are 3600rpm?!
6
32
→ More replies (6)13
103
u/Different_Fix_2217 20d ago
39
u/ChristRedeemsSinners 20d ago
Weird that non-thinking support is an API only feature.
9
u/fantasticsid 20d ago
Based on my experience with 3.6, prefilling
<think>\n\n</think>\n- like the various jinja templates do - to disable thinking works probably 90-95% of the time. The other ~5-10% of the time, the model thinks anyway and emits a second</think>when it's done. It's possible that 3.8 has the same behaviour and the official API has some way of detecting/working around this that would look pretty damn stupid if they released it. If you look at the 3.8 jinja template, the "reasoning effort" isn't implemented terribly cleverly - it just talks to the model in the second person and asks it to reason less.Reasoning control has always been a weakness of the Qwen models, so this doesn't surprise me.
Lack of mmproj, however, feels like a deliberate attempt at market segmentation. Given that those just decode image data into tokens, and the whole Qwen family shares a vocab, I do wonder if it'd be possible to hack the mmproj from some other Qwen model into use here, though.
→ More replies (1)3
u/hellomistershifty 20d ago
I love how AI development is a mix of wild cutting edge research, weird hacks, and asking it nicely to behave
48
u/Reactor-Licker 20d ago
That’s weird, why did they remove vision? Hopefully Qwen 3.8 27B doesn’t do the same thing.
18
17
u/vincentz42 20d ago
Meanwhile the Qwen3.8 27B open weight that is due in two days does have vision. This has to be intentional, right?
→ More replies (3)→ More replies (1)12
u/Hoak-em 20d ago
Wait, vision is what made this model good — without it it’s just a big, expensive to run somewhat smart llm
→ More replies (2)
383
u/Legal-Ad-3901 20d ago
5tb bf16 jfc. even the crazy home lab kids cant hang anymore
158
u/hyperrealists 20d ago
Poor me can’t even download it lol
71
u/EndlessZone123 20d ago
raid 0 hdd time.
12
u/LukeLikesReddit 20d ago
I mean i have 16tb ready, I can download it, will I be able to do anything with it? Absolutely not lol.
6
u/Ell2509 20d ago
Same lol. I have saved glm 5.2, DS4, MM3, and even kimi k3. Can't use it, but i have it now haha.
3
u/LukeLikesReddit 20d ago
haha yeah agreed, I just asked a friend if I could borrow their server blade to run this :) Their IT systems dont need it that badly.
→ More replies (2)6
u/Koakie 20d ago
The digital equivalent of a glorified paperweight
5
u/LukeLikesReddit 20d ago
I like to call myself a historian. Though by the time I download it we will be on something else XD
17
u/ChristRedeemsSinners 20d ago
Lol, that's what I was thinking. I need to spend 1k just to download it.
→ More replies (2)22
u/jikilan_ 20d ago
Don’t need to download the full copy , you can stream it.
One of llama.cpp PR support streaming from disk and even cloud storage if I remember correctly 😘
→ More replies (1)109
u/xPXpanD llama.cpp 20d ago
Years/token is my favorite metric.
30
u/Think_Wing_1357 20d ago
After a few million years, you may get 42
12
u/vivekkhera 20d ago
Then you have to build a new server just to figure out what the question was.
6
3
41
u/CapeChill 20d ago
I can't even run it at work... Crazy we're passing what the 8xH100 boxes can even do on open models now. You'd need multiple racks of H100s for that.
6
u/Own_Anything9292 20d ago
looks like 16xh200s tp 16 can run it according to vllm, so 4 node h100s tp 32 might be the trick. NVFP4 published instead of NVFP8 means we need blackwell and not hopper :( you’ll probably need to hack something together to make nvfp4 work with h100s
4
u/CapeChill 20d ago
You can do 4 air cooled nodes reasonably in a extra tall rack I guess. It's also crazy to see that six figure boxes can't run the latest model encoding.
→ More replies (1)→ More replies (1)14
u/chithanh 20d ago
I guess it is the ultimate troll, release models that are so large that you can run them locally on Chinese hardware only, because running them on NVIDIA hardware would bankrupt you
Reports are that 5T and 10T models are being prepared
→ More replies (1)7
u/CapeChill 20d ago
I work in HPC so this doesn't really land. I get to tinker with a few H100s because companies are happy to drop a few million to run a open weight model in house with highly confidential data on.
I love my little home weather network and the prediction it does and its cute what my local hardware can compute. I've installed HPC clusters that do climate simulation, that will never run well locally on current hardware and that's okay. Same for 5-10t open weight models, it will hopefully continue to be the case there are open weights so huge only research can justify running them as it attracts brainiacs that will actually trickle down to us peons.
Source: sometimes I get to be a fly on the wall when these academic brainiacs talk.29
u/Jolly_Criticism9190 20d ago
As someone who just bought 256GB of ram. I concur. Holy moly
19
u/rinmperdinck 20d ago
Why didn't you buy 5TB of RAM instead?
9
u/Jolly_Criticism9190 20d ago
Some guy named Sam flied to Korean ahead of me before I could convince a guy named Jensen to leave out some DRAM capacity for all human kind
→ More replies (1)3
19
u/quantgorithm 20d ago
Guess I'm waiting for the 27B.
→ More replies (1)4
u/-dysangel- 20d ago
at least with the 3.5 series for coding, 27B seemed better than the larger models anyway
20
u/DocMadCow 20d ago
Bold to assume there isn't a homelab person out there that doesn't have this much ram :D
39
u/bruns20 20d ago
Thats not a home anymore lmao, thats just a personal data centre
12
u/thejinx0r 20d ago
7
u/DocMadCow 20d ago
Exactly this bro hasn't spent any time in those eccentric Subs. Every time I go there I realize not only can I not afford the hardware but especially not the electrical bill.
→ More replies (1)9
u/Legal-Ad-3901 20d ago
I mean I have 1.5tb ddr4 and 944gb vram so crawl speeds for decent quant are in grasp. But useable? Definitely feels like a line in the sand is happening on intelligence for the proliteriate
26
u/Makers7886 20d ago
it's makes my 12x3090 + 256gb epyc machine feel like the guy trying to run qwen 9b on his igpu and streaming weights from a hard drive making clicking noises.
→ More replies (3)6
4
u/FullstackSensei llama.cpp 20d ago
I searched in vein for any mention of data types in the model card and blog post. 200+ files in alternating 17 and 34GB each.
I'm reworking my homelab to be able to run full K3 across 2 machines, but this one is just too much, especially in light of K3 and the just announced DS4 pro (which I can run on a single machine).
3
u/allenasm 20d ago
what? you mean with my mac m3 studio 512gb unified and 8tb pcie 5.0 nvme drive? 'hold my beer'... :)
→ More replies (2)→ More replies (3)3
u/_TheWolfOfWalmart_ 20d ago
Here I was thinking I'm so cool a 256 GB VRAM server.
The box has 768 GB system RAM, I could run the UD-IQ1_S with offloading, but fuck... that's not going to be a fun experience.
67
u/SandySkittle 20d ago
Now burn this to a chip so we can run it at 16k tokens per second and we’re good.
45
u/SmartCustard9944 20d ago
You need a silicon slab that is 0.7 meters in diameter by the way
36
→ More replies (2)30
u/Boreras 20d ago
Damn I just checked, the Llama 3.1 8B chip is fucking 815 mm², 53 billion transistors. (on the older tsmc 6nm process)
Same as H100. This is basically the limit for litho machines, called the reticle limit. This is the limit of the photomask (basically the negative of the chip which the EUV patterns on the photo resist).
Taalas literally cannot build anything larger than 40b parameters today, although I assume 2NP yields do not permit chips at the reticle limit.
→ More replies (5)→ More replies (4)7
u/Admirable_Market2759 20d ago edited 20d ago
If it was cheaper than buying GPUs, then I’d buy an ASIC just for K3.
→ More replies (3)
123
u/Piyh 20d ago
95B active is cray. Scaling gonna scale.
41
u/fgk55555 20d ago
My entire rig could run one of the experts (quantized) at maybe 1tkps.
9
u/RegisteredJustToSay 20d ago
Look at fancy pants money bags over here. Pretty sure mine would spontaneously combust at the mere suggestion.
27
u/SandySkittle 20d ago
Yes crazy, but frankly I think some MoE go too far with low active numbers. Or rather, I would really like a 122b-30a model that still fits in midrange enthusiast local setups (4x r9700 ) has a lot of world knowledge (more than a 30b model) but dares to keep the active number on the higher end to preserve more of the qualities of a dense model.
→ More replies (1)5
u/Carbonite1 20d ago
I had this same opinion for a while but I've been starting to come around a bit -- I mean, even mid-sized models are so sparse these days, like DSV4F being >200B but only 13B active, and are seeing such good results -- I can only imagine the labs have tried a higher proportion of active parameters and it isn't even close to worth the tradeoff?
3
u/SandySkittle 20d ago
It depends on the task. For very complex multi facetted analysis with many components interacting with each other in a nuanced way with a lot of nuanced context (context not per se kv context, but in the general meaning of the word), that require large very detailed structured prompts with lots of caveats to frame the question, like complex legal analysis that can only be partially broken down, 13b active parameters, even with max/deep (but still sequential!) reasoning is just missing the depth. It starts to lose or compress details in its reasoning or misses connections between details. You get junior analyst answers to senior analist questions. This is why larger dense models (70b plus) are so important.
There is more to llms than how well they do coding..
That’s why I would like to see larger but still locally feasible models (up to 160 gb at q8) with a lot of world knowledge but with larger active parameters than just 13b.
3
u/FullstackSensei llama.cpp 20d ago
K3 is 100B+ active, but they SFT'd the whole thing in fp4. The routed experts are 33GB per token. A ton for sure, but doable on a dual DDR4 Xeon or Epyc.
This is fp16 through and through. The full fat is ~5TB, active possibly ~160GB.
Sure, you can quantize, but that will inadvertently reduce intelligence.a
8
u/Maleficent-Ad5999 20d ago
I wish someone comes up with Mixtures of MOEs
→ More replies (2)6
u/Badger-Purple 20d ago
You can do this with an agent harness. Hermes supports a Mixture of Agents mode where you collate several LLM answers and use a main model to ingest them and synthesize a final answer.
→ More replies (2)3
3
480
u/ApprehensiveTart3158 20d ago
Finally a model I can run locally, took them so long to release a model at a reasonable size
168
u/ScreenAppropriate679 20d ago
I run it in my local datacenter no problem
103
u/ApprehensiveTart3158 20d ago
Just connect a 6TB hard drive to your raspberry pi, it will run it! (maybe)
54
u/Asleep_Document9811 20d ago
I build an inferencing engine using a small African village as a substrate. Funding pleeeeeease!
31
u/ApprehensiveTart3158 20d ago
Why do that? It's math, you can calculate llm matrices on paper
14
u/Vegetable-Clerk9075 20d ago
At how many tokens per week?
→ More replies (1)6
→ More replies (1)18
u/libregrape llama.cpp 20d ago
On paper?! Why waste paper and pens when everyone knows how to multiply tensors in head!
12
u/Strawberry3141592 20d ago
Nah, what you wanna do is get several billion TI-84 calculators (3MB storage each), and network them all together into the world's most fuckass cluster. 1 token per day, maybe.
11
4
→ More replies (4)3
u/AvengerDr 20d ago
The Three Body Problem way: get a hundred thousand or preferably million people in a field. Have each of them hold a flag and tell them to raise it if they are a 1 or keep it lowered if they are a 0.
Then multiple dudes on horses just run down the lines and give them instructions.
4
→ More replies (3)37
u/Real_Ebb_7417 20d ago
Posts "I made Qwen3.8 Max run on my toaster with this new technique" over the next month incoming.
(disclaimer: the "new" technique is streaming from SSD and Qwen runs at 0.01 tps)
(disclaimer 2: half the comments will be "It will damage your ssd" and the other half "llama.cpp has been handling this for the long time already")
(disclaimer 3: The responses to the first half comments will be "it won't damage your ssd")
→ More replies (4)
38
u/nickm_27 llama.cpp 20d ago
Customizable reasoning effort is a nice improvement.
https://huggingface.co/Qwen/Qwen3.8-2.4T-A95B#qwen38-highlights
→ More replies (2)
56
u/Technical-Earth-3254 20d ago
Do I read it correctly that the open weight version has no vision support?
50
u/Fristender 20d ago
From https://modelscope.cn/models/Qwen/Qwen3.8-2.4T-A95B :
In particular, Qwen3.8-Max is the official version based on Qwen3.8-2.4T-A95B with more features, such as vision input & non-thinking support, 1M context length by default, official built-in tools, etc.
30
u/Technical-Earth-3254 20d ago
Yeah, that's why I asked. That's the same text as on hf. This kinda confuses me, why they remove vision when all the other big models now come with it
15
u/Song-Historical 20d ago
It's a specialization that you need to train separately. It makes sense, you spend your money training a purely text based model, for other teams to adapt as needed to vision etc
→ More replies (1)3
→ More replies (2)18
u/wren6991 20d ago
Gonna make a wild prediction: the model itself is still vision-trained, and we will figure out how to glue one of the existing Qwen vision adapters to it.
I think this is a profitability concession to keep their leadership happy. I don't see them training two different versions of a 2.4T model just for the sake of an open-weight release.
27
29
u/hebelehubele 20d ago
I have found a qwen3824T.exe on internet, can i just run it? It is only 2.4kB. Yay…
30
u/YOMUMSOBIG 20d ago
Finally! Perfect size for running it on my smart watch.
8
u/SolenoidSoldier 20d ago
Real talk...is there a conversational AI that can run on today's smart watches? That would be sweet
→ More replies (1)
27
14
u/TheRealMasonMac 20d ago
Like the rumors suggested, they are also opting for a revenue share model like MoonshotAI and MiniMax. Seems like this is the direction open-weight releases in China are going.
12
10
10
9
u/ideaofsoul 20d ago
Great! I just need another 99 rtx 3090 and its ready to serve
→ More replies (2)
9
8
15
6
8
8
u/Iterative_One 20d ago
Do I have enough?
VRAM - No
RAM - Nope
Hard Drive Capacity - Absolutely Not.
6
u/TheLocalLab 20d ago
Unsloth GGUF Quants was also available - https://huggingface.co/unsloth/Qwen3.8-2.4T-A95B-GGUF
18
u/Embarrassed_Adagio28 20d ago
After a week of using this model, my disappointment is immeasurable and my day is ruined /s
5
u/Medium_Chemist_4032 20d ago
95b is exceptionally deep for bigger MoE's, holy smokes... This might compete
6
5
4
12
3
4
7
u/RickyRickC137 20d ago edited 20d ago
After loading this model, I’m not sure what I’m going to do with all the VRAM I have left.
→ More replies (2)
3
3
u/Hefty_Wolverine_553 20d ago
Reasoning Content: Set the maximum output length to 262,144 tokens.
Final Response: Set the maximum output length to 131,072 tokens.
Using up the entirety of Qwen3.6 27B's context window for reasoning alone...
3
u/exaknight21 20d ago
I need me a 1 bit awq, that is then 1 bit’d again. Or a 4 bit that is 1 bit’d.
→ More replies (2)
3
u/fooo12gh 20d ago
If the model parameters increase at such a rate, I doubt we'll see RAM prices dropping down anytime soon.
3
8
u/JsThiago5 20d ago
Why hype this when 99.999% of people cannot run it? I was really expecting 27b today :(
13
u/banana_slurp_jug 20d ago
Firstly, 27b is in less than 48 hours. Secondly, even if you can't run an open-weights model on your own computer, the prices to pay for inference with it is cheaper since multiple providers will compete for value.
→ More replies (7)3
u/GregAbeI 20d ago
I think way fewer than 1 in 100,000 people can run this, which is what 99.999% means.
99.999999% of people can't run this.
4
u/amy-schumer-tampon 20d ago
I have the feeling that a 2.4T model isn't something many people can run locally.
2
2
u/ocean_protocol 20d ago
this is wild, 95B active out of 2.4T total is a pretty aggressive sparsity ratio. anyone know what the expert routing setup looks like on this one? And how it compares to deepseek's approach
2
2
u/madjesta 20d ago
Could this even be run offline in any practical way just to distill? Maybe per layer batching? JFC this is huge.
2
u/Then_Blueberry7290 20d ago
text only? No vision capabilities? Or just the unsloth version doesn't contain? AFAIK qwen 3.8 27b will be vision enabled model, or not?
Anyway q1 is 400GB size...
2
u/Legitimate-Dog5690 20d ago
Amazing stuff, so glad to have this at home. Only need 5x RTX 6000s and I can run Q1.
→ More replies (2)
2
2
u/xXDennisXx3000 20d ago
When Kimi K3 was a tank, Qwen 3.8 Max is a spaceship....
→ More replies (1)
2
u/Beltalowdamon 20d ago
Guessing it will be a while until you can run a distilled version of this on 8gb vram and 32gb ram!
2





262
u/No_War_8891 20d ago
I can run the active part locally lol