r/LocalLLaMA 20d ago

New Model Qwen3.8-2.4T-A95B Released

https://huggingface.co/Qwen/Qwen3.8-2.4T-A95B
1.6k Upvotes

397 comments sorted by

262

u/No_War_8891 20d ago

I can run the active part locally lol

27

u/Sea-Ad-5390 20d ago

I can help host some of the offloaded experts for you streamed across the internet

16

u/No_War_8891 20d ago

I read people really did that - degens I love it

→ More replies (3)

10

u/BobbyL2k 20d ago

I wonder how good it would be if we get rid of MoE router and just lock to a specific set of experts. Make it a 95B dense model.

12

u/Jump3r97 20d ago

Interesting question but I think it would be abysmal. The Expert selected can Change wildly per individual token

6

u/[deleted] 20d ago

[deleted]

→ More replies (1)
→ More replies (1)
→ More replies (5)

166

u/lm-gtfy 20d ago

Finally. What is the meaning of life?

38

u/xaeru 20d ago

27

u/tripazardly 20d ago

Oh no, not again.

11

u/Caffdy 20d ago

OOTL, what is that?

17

u/corysama 20d ago

42 is famously a joke from Hitchhiker's Guide to the Galaxy. A bowl of petunias is less famously the same.

HHGttG if full of jokes that build up over long stretches (multiple books sometimes) and punchline in a single sentence. Can't recommend enough.

3

u/Caffdy 20d ago

so the petunias are from the same book?

→ More replies (2)

11

u/Paganator 20d ago

Many people have speculated that if we knew exactly why the bowl of petunias had thought that we would know a lot more about the nature of the universe than we do now.

→ More replies (1)

6

u/feelspeaceman 20d ago

It's still good for distilling to train smaller models to be smarter, basically the popular 27B was distilled from this giant 2.4T.

A good mother model will likely give birth to more capable smaller models, but it's still requiring a tons of trials and errors, those that costs money to retry.

3

u/JahJedi 20d ago

Personal cluster that can ran it for all the femily

→ More replies (4)

217

u/Intelligent_Ice_113 20d ago edited 20d ago

what is the knowledge cutoff date?

299

u/Super_Range45 20d ago

yes

129

u/hyperrealists 20d ago

Thank god

18

u/srigi 20d ago

"But wait, ..."

→ More replies (1)

24

u/HungryMachines 20d ago

I think you are off by a few

12

u/BigBrainGoldfish 20d ago

So unhelpful but still so perfect. Lol

→ More replies (1)

91

u/JumpingJack79 20d ago

I kinda like early cutoff dates. Whenever a model tells me its cutoff is 2024, I'm like "Dude, it's 2026 now and you won't believe what's happened..." 😏

129

u/-dysangel- 20d ago

"The user is talking about a hypothetical timeline"

21

u/JumpingJack79 20d ago

Lol, yes. And to be fair, I'd be skeptical too if I was in their place hearing those things.

20

u/this_is_a_long_nickn 20d ago

“I hate when users hallucinate”

3

u/MoodDelicious3920 20d ago

Agi 😂 when model starts laughing on users!

→ More replies (1)

19

u/Mickenfox 20d ago

In a few decades people are going to be doing 2020s roleplay with ancient models.

2

u/that_one_guy63 20d ago

Would be interesting to train on only text before 1970. Or like way earlier. It would be so confused.

→ More replies (2)

6

u/AvidCyclist250 llama.cpp 20d ago

The user seems to be talking about a fictional future scenario

3

u/JumpingJack79 20d ago

I've heard this one too. Oh, I wish I were! 😭

44

u/teleprax 20d ago edited 20d ago

One of my favorite LLM memories was telling GPT (5 maybe?) a bunch of stuff like elon doing the nazi salute and some other stuff going on around that time, and then it proceeds to lecture me on fake news and how none of that is remotely correct and how it would be super serious.

Then I tell it to do a web search and it starts it's next reply with:

"Ok. wow."


EDIT: Ok i found it, its not quite the same as I remember, but close enough

https://i.imgur.com/Ta33JIT.jpeg https://i.imgur.com/kPDa9tm.jpeg

25

u/personahorrible 20d ago

"Sounds like a bizarre Black Mirror meets Idiocracy arc."

Bruh. You have no idea.

15

u/Valuable_Cow2596 20d ago

Oh boy that gave me a chuckle. Thanks for sharing. 

14

u/OttoRenner 20d ago

Mine went bonkers after I showed it several screenshots and links of GPU prices...was very funny to witness 😂😂

11

u/AlpacaDC 20d ago

“You’re right” lmao

→ More replies (1)

45

u/Ok_Ocelot2268 20d ago

Cutoff: April 2026. Startoff: June 5, 1989

14

u/thepaligator 20d ago

I remember June 5th. Not a lot of foot traffic that day. Clear Streets. It was nice.

10

u/This-Consequence-957 20d ago

Bad Boy, I know because my birthdate is June 4 🙈

8

u/goldrunout 20d ago

What is a start off date? Do they not train on older sources?

23

u/Anwar6969 20d ago

that’s a tiananmen square joke. nothing happened on june 4th 1989

21

u/FastDecode1 20d ago

[ Removed by the CCP ]

12

u/LawfulLeah 20d ago

the fact that the comment you replied to was removed by reddit makes this 10x funnier

6

u/Anwar6969 20d ago

crazy shit, i had to appeal to lift off the warning and the comment ban. i apparently got flagged for rule 1

→ More replies (1)

5

u/c_glib 20d ago

Is a joke on a Chinese model. Remember June 4th 1989?

97

u/FullstackSensei llama.cpp 20d ago

Because, that's the only thing holding you from running a 5TB model?

25

u/fullup72 20d ago

You can always quantize to 1 bit and stream experts from a 5400rpm HDD.

12

u/FullstackSensei llama.cpp 20d ago

Why settle for 5400rpm when you can get old drives that are 3600rpm?!

6

u/HulksInvinciblePants 20d ago

Stack em for 8000rpm throughput

32

u/MrObsidian_ 20d ago

Probably tomorrow

13

u/jikilan_ 20d ago

It know what you did in the last summer

→ More replies (6)

103

u/Different_Fix_2217 20d ago

Be warned they state its not the same capabilities as the full API version. Such as not having vison.

39

u/ChristRedeemsSinners 20d ago

Weird that non-thinking support is an API only feature.

9

u/fantasticsid 20d ago

Based on my experience with 3.6, prefilling <think>\n\n</think>\n - like the various jinja templates do - to disable thinking works probably 90-95% of the time. The other ~5-10% of the time, the model thinks anyway and emits a second </think> when it's done. It's possible that 3.8 has the same behaviour and the official API has some way of detecting/working around this that would look pretty damn stupid if they released it. If you look at the 3.8 jinja template, the "reasoning effort" isn't implemented terribly cleverly - it just talks to the model in the second person and asks it to reason less.

Reasoning control has always been a weakness of the Qwen models, so this doesn't surprise me.

Lack of mmproj, however, feels like a deliberate attempt at market segmentation. Given that those just decode image data into tokens, and the whole Qwen family shares a vocab, I do wonder if it'd be possible to hack the mmproj from some other Qwen model into use here, though.

3

u/hellomistershifty 20d ago

I love how AI development is a mix of wild cutting edge research, weird hacks, and asking it nicely to behave

→ More replies (1)

48

u/Reactor-Licker 20d ago

That’s weird, why did they remove vision? Hopefully Qwen 3.8 27B doesn’t do the same thing.

18

u/vincentz42 20d ago

According to the signup page it will.

17

u/vincentz42 20d ago

Meanwhile the Qwen3.8 27B open weight that is due in two days does have vision. This has to be intentional, right?

→ More replies (3)

12

u/Hoak-em 20d ago

Wait, vision is what made this model good — without it it’s just a big, expensive to run somewhat smart llm

→ More replies (2)
→ More replies (1)

383

u/Legal-Ad-3901 20d ago

5tb bf16 jfc. even the crazy home lab kids cant hang anymore

158

u/hyperrealists 20d ago

Poor me can’t even download it lol

71

u/EndlessZone123 20d ago

raid 0 hdd time.

12

u/LukeLikesReddit 20d ago

I mean i have 16tb ready, I can download it, will I be able to do anything with it? Absolutely not lol.

6

u/Ell2509 20d ago

Same lol. I have saved glm 5.2, DS4, MM3, and even kimi k3. Can't use it, but i have it now haha.

3

u/LukeLikesReddit 20d ago

haha yeah agreed, I just asked a friend if I could borrow their server blade to run this :) Their IT systems dont need it that badly.

6

u/Koakie 20d ago

The digital equivalent of a glorified paperweight

5

u/LukeLikesReddit 20d ago

I like to call myself a historian. Though by the time I download it we will be on something else XD

→ More replies (2)

17

u/ChristRedeemsSinners 20d ago

Lol, that's what I was thinking. I need to spend 1k just to download it.

22

u/jikilan_ 20d ago

Don’t need to download the full copy , you can stream it.

One of llama.cpp PR support streaming from disk and even cloud storage if I remember correctly 😘

109

u/xPXpanD llama.cpp 20d ago

Years/token is my favorite metric.

30

u/Think_Wing_1357 20d ago

After a few million years, you may get 42

12

u/vivekkhera 20d ago

Then you have to build a new server just to figure out what the question was.

6

u/techno156 20d ago

And then someone decides to blow it up for a highway.

3

u/hyperrealists 20d ago

YTFT is insane I hear.

→ More replies (1)
→ More replies (2)

41

u/CapeChill 20d ago

I can't even run it at work... Crazy we're passing what the 8xH100 boxes can even do on open models now. You'd need multiple racks of H100s for that.

6

u/Own_Anything9292 20d ago

looks like 16xh200s tp 16 can run it according to vllm, so 4 node h100s tp 32 might be the trick. NVFP4 published instead of NVFP8 means we need blackwell and not hopper :( you’ll probably need to hack something together to make nvfp4 work with h100s

4

u/CapeChill 20d ago

You can do 4 air cooled nodes reasonably in a extra tall rack I guess. It's also crazy to see that six figure boxes can't run the latest model encoding.

→ More replies (1)

14

u/chithanh 20d ago

I guess it is the ultimate troll, release models that are so large that you can run them locally on Chinese hardware only, because running them on NVIDIA hardware would bankrupt you

Reports are that 5T and 10T models are being prepared

7

u/CapeChill 20d ago

I work in HPC so this doesn't really land. I get to tinker with a few H100s because companies are happy to drop a few million to run a open weight model in house with highly confidential data on.

I love my little home weather network and the prediction it does and its cute what my local hardware can compute. I've installed HPC clusters that do climate simulation, that will never run well locally on current hardware and that's okay. Same for 5-10t open weight models, it will hopefully continue to be the case there are open weights so huge only research can justify running them as it attracts brainiacs that will actually trickle down to us peons.
Source: sometimes I get to be a fly on the wall when these academic brainiacs talk.

→ More replies (1)
→ More replies (1)

29

u/Jolly_Criticism9190 20d ago

As someone who just bought 256GB of ram. I concur. Holy moly

19

u/rinmperdinck 20d ago

Why didn't you buy 5TB of RAM instead?

9

u/Jolly_Criticism9190 20d ago

Some guy named Sam flied to Korean ahead of me before I could convince a guy named Jensen to leave out some DRAM capacity for all human kind

3

u/rinmperdinck 20d ago

Wow those guys sound like pricks

→ More replies (1)

19

u/quantgorithm 20d ago

Guess I'm waiting for the 27B.

4

u/-dysangel- 20d ago

at least with the 3.5 series for coding, 27B seemed better than the larger models anyway

→ More replies (1)

20

u/DocMadCow 20d ago

Bold to assume there isn't a homelab person out there that doesn't have this much ram :D

39

u/bruns20 20d ago

Thats not a home anymore lmao, thats just a personal data centre

12

u/thejinx0r 20d ago

7

u/DocMadCow 20d ago

Exactly this bro hasn't spent any time in those eccentric Subs. Every time I go there I realize not only can I not afford the hardware but especially not the electrical bill.

→ More replies (1)

9

u/Legal-Ad-3901 20d ago

I mean I have 1.5tb ddr4 and 944gb vram so crawl speeds for decent quant are in grasp. But useable? Definitely feels like a line in the sand is happening on intelligence for the proliteriate

26

u/Makers7886 20d ago

it's makes my 12x3090 + 256gb epyc machine feel like the guy trying to run qwen 9b on his igpu and streaming weights from a hard drive making clicking noises.

6

u/tripplebeamteam 20d ago

Hey that’s me, running qwen 35B A3B on my igpu and getting 4 tok/sec!

→ More replies (3)

4

u/FullstackSensei llama.cpp 20d ago

I searched in vein for any mention of data types in the model card and blog post. 200+ files in alternating 17 and 34GB each.

I'm reworking my homelab to be able to run full K3 across 2 machines, but this one is just too much, especially in light of K3 and the just announced DS4 pro (which I can run on a single machine).

3

u/allenasm 20d ago

what? you mean with my mac m3 studio 512gb unified and 8tb pcie 5.0 nvme drive? 'hold my beer'... :)

→ More replies (2)

3

u/_TheWolfOfWalmart_ 20d ago

Here I was thinking I'm so cool a 256 GB VRAM server.

The box has 768 GB system RAM, I could run the UD-IQ1_S with offloading, but fuck... that's not going to be a fun experience.

→ More replies (3)

67

u/SandySkittle 20d ago

Now burn this to a chip so we can run it at 16k tokens per second and we’re good.

45

u/SmartCustard9944 20d ago

You need a silicon slab that is 0.7 meters in diameter by the way

36

u/SandySkittle 20d ago

I have room in the attic

30

u/Boreras 20d ago

Damn I just checked, the Llama 3.1 8B chip is fucking 815 mm², 53 billion transistors. (on the older tsmc 6nm process)

https://taalas.com/products/

Same as H100. This is basically the limit for litho machines, called the reticle limit. This is the limit of the photomask (basically the negative of the chip which the EUV patterns on the photo resist).

Taalas literally cannot build anything larger than 40b parameters today, although I assume 2NP yields do not permit chips at the reticle limit.

→ More replies (5)
→ More replies (2)

7

u/Admirable_Market2759 20d ago edited 20d ago

If it was cheaper than buying GPUs, then I’d buy an ASIC just for K3.

→ More replies (3)
→ More replies (4)

123

u/Piyh 20d ago

95B active is cray. Scaling gonna scale.

41

u/fgk55555 20d ago

My entire rig could run one of the experts (quantized) at maybe 1tkps.

9

u/RegisteredJustToSay 20d ago

Look at fancy pants money bags over here. Pretty sure mine would spontaneously combust at the mere suggestion.

27

u/SandySkittle 20d ago

Yes crazy, but frankly I think some MoE go too far with low active numbers. Or rather, I would really like a 122b-30a model that still fits in midrange enthusiast local setups (4x r9700 ) has a lot of world knowledge (more than a 30b model) but dares to keep the active number on the higher end to preserve more of the qualities of a dense model.

5

u/Carbonite1 20d ago

I had this same opinion for a while but I've been starting to come around a bit -- I mean, even mid-sized models are so sparse these days, like DSV4F being >200B but only 13B active, and are seeing such good results -- I can only imagine the labs have tried a higher proportion of active parameters and it isn't even close to worth the tradeoff?

3

u/SandySkittle 20d ago

It depends on the task. For very complex multi facetted analysis with many components interacting with each other in a nuanced way with a lot of nuanced context (context not per se kv context, but in the general meaning of the word), that require large very detailed structured prompts with lots of caveats to frame the question, like complex legal analysis that can only be partially broken down, 13b active parameters, even with max/deep (but still sequential!) reasoning is just missing the depth. It starts to lose or compress details in its reasoning or misses connections between details. You get junior analyst answers to senior analist questions. This is why larger dense models (70b plus) are so important.

There is more to llms than how well they do coding..

That’s why I would like to see larger but still locally feasible models (up to 160 gb at q8) with a lot of world knowledge but with larger active parameters than just 13b.

→ More replies (1)

3

u/FullstackSensei llama.cpp 20d ago

K3 is 100B+ active, but they SFT'd the whole thing in fp4. The routed experts are 33GB per token. A ton for sure, but doable on a dual DDR4 Xeon or Epyc.

This is fp16 through and through. The full fat is ~5TB, active possibly ~160GB.

Sure, you can quantize, but that will inadvertently reduce intelligence.a

8

u/Maleficent-Ad5999 20d ago

I wish someone comes up with Mixtures of MOEs

6

u/Badger-Purple 20d ago

You can do this with an agent harness. Hermes supports a Mixture of Agents mode where you collate several LLM answers and use a main model to ingest them and synthesize a final answer.

→ More replies (2)
→ More replies (2)

3

u/SmartCustard9944 20d ago

What if we MoE the active params too 🤔

3

u/fluffysheap 20d ago

A Cray is what you need to run it

480

u/ApprehensiveTart3158 20d ago

Finally a model I can run locally, took them so long to release a model at a reasonable size

168

u/ScreenAppropriate679 20d ago

I run it in my local datacenter no problem

103

u/ApprehensiveTart3158 20d ago

Just connect a 6TB hard drive to your raspberry pi, it will run it! (maybe)

54

u/Asleep_Document9811 20d ago

I build an inferencing engine using a small African village as a substrate. Funding pleeeeeease!

31

u/ApprehensiveTart3158 20d ago

Why do that? It's math, you can calculate llm matrices on paper

14

u/Vegetable-Clerk9075 20d ago

At how many tokens per week?

6

u/ApprehensiveTart3158 20d ago

Matters how fast you can calculate

8

u/MaruluVR 20d ago

This is one of those rare cases where more training will increase the tp/s

→ More replies (1)

18

u/libregrape llama.cpp 20d ago

On paper?! Why waste paper and pens when everyone knows how to multiply tensors in head!

18

u/xaeru 20d ago

Why in your head? Why waste precious brain cells? Just hold two magnetized rocks, close your eyes, and let quantum fluctuations handle the attention weights.

3

u/MmmmMorphine 20d ago

Yeah but that's super lossy. I recommend using two magnetized toddlers.

→ More replies (1)

12

u/Strawberry3141592 20d ago

Nah, what you wanna do is get several billion TI-84 calculators (3MB storage each), and network them all together into the world's most fuckass cluster. 1 token per day, maybe.

11

u/x10der_by 20d ago

1 token per day

3

u/Terminator857 20d ago

Weren't computers so fast 25 years ago?

4

u/ApeGrower 20d ago

With 35 tok/year!

3

u/AvengerDr 20d ago

The Three Body Problem way: get a hundred thousand or preferably million people in a field. Have each of them hold a flag and tell them to raise it if they are a 1 or keep it lowered if they are a 0.

Then multiple dudes on horses just run down the lines and give them instructions.

→ More replies (4)

4

u/volleyneo 20d ago

Is this why the Danube River has severe draught? It was you bastards!

37

u/Real_Ebb_7417 20d ago

Posts "I made Qwen3.8 Max run on my toaster with this new technique" over the next month incoming.

(disclaimer: the "new" technique is streaming from SSD and Qwen runs at 0.01 tps)

(disclaimer 2: half the comments will be "It will damage your ssd" and the other half "llama.cpp has been handling this for the long time already")

(disclaimer 3: The responses to the first half comments will be "it won't damage your ssd")

8

u/eidrag 20d ago

120s/tok!

→ More replies (4)

3

u/Mr-I17 20d ago

I'm pretty confident that I can run it locally at 60 tokens per hour.

→ More replies (3)

38

u/nickm_27 llama.cpp 20d ago

Customizable reasoning effort is a nice improvement.

https://huggingface.co/Qwen/Qwen3.8-2.4T-A95B#qwen38-highlights

→ More replies (2)

56

u/Technical-Earth-3254 20d ago

Do I read it correctly that the open weight version has no vision support?

50

u/Fristender 20d ago

From https://modelscope.cn/models/Qwen/Qwen3.8-2.4T-A95B :

In particular, Qwen3.8-Max is the official version based on Qwen3.8-2.4T-A95B with more features, such as vision input & non-thinking support, 1M context length by default, official built-in tools, etc.

30

u/Technical-Earth-3254 20d ago

Yeah, that's why I asked. That's the same text as on hf. This kinda confuses me, why they remove vision when all the other big models now come with it

15

u/Song-Historical 20d ago

It's a specialization that you need to train separately. It makes sense, you spend your money training a purely text based model, for other teams to adapt as needed to vision etc

3

u/Balance- 20d ago

Interesting. So this is kind of a stripped down / base model?

→ More replies (1)
→ More replies (1)

18

u/wren6991 20d ago

Gonna make a wild prediction: the model itself is still vision-trained, and we will figure out how to glue one of the existing Qwen vision adapters to it.

I think this is a profitability concession to keep their leadership happy. I don't see them training two different versions of a 2.4T model just for the sake of an open-weight release.

→ More replies (2)

27

u/SoupDue6629 20d ago

Somebody REAP it to 35B pls k thx

→ More replies (1)

29

u/hebelehubele 20d ago

I have found a qwen3824T.exe on internet, can i just run it? It is only 2.4kB. Yay…

11

u/andr386 20d ago

That's fine. But you should make sure to download enough VRAM before running it.

30

u/YOMUMSOBIG 20d ago

Finally! Perfect size for running it on my smart watch.

8

u/SolenoidSoldier 20d ago

Real talk...is there a conversational AI that can run on today's smart watches? That would be sweet

→ More replies (1)

27

u/Septerium 20d ago

Finally we have a good successor for Qwen 3.5 9B as a local daily driver

14

u/TheRealMasonMac 20d ago

Like the rumors suggested, they are also opting for a revenue share model like MoonshotAI and MiniMax. Seems like this is the direction open-weight releases in China are going.

12

u/milkipedia 20d ago

HuggingFace gonna crash today

11

u/d70 20d ago

Fit nicely on my GTX 980

10

u/Automatic-Boot665 20d ago

Can’t wait to run it at q0.1

9

u/ideaofsoul 20d ago

Great! I just need another 99 rtx 3090 and its ready to serve

→ More replies (2)

9

u/Feztopia 20d ago

0.01 bit quant when?

→ More replies (2)

8

u/Daniel_H212 20d ago

Vision encoder not released, will it be coming later?

15

u/mxforest 20d ago

One RTX Pro 6000 per expert Quantized. Holy balls.

→ More replies (1)

6

u/MrVeinless 20d ago

I am surprised it’s not multimodal even with an mmproj.

8

u/AlternateWitness 20d ago

Can anyone lend me some ram?

→ More replies (1)

8

u/Iterative_One 20d ago

Do I have enough?

VRAM - No

RAM - Nope

Hard Drive Capacity - Absolutely Not.

18

u/Embarrassed_Adagio28 20d ago

After a week of using this model, my disappointment is immeasurable and my day is ruined /s

5

u/Medium_Chemist_4032 20d ago

95b is exceptionally deep for bigger MoE's, holy smokes... This might compete

6

u/SandySkittle 20d ago

Yes, 122b a30+ please :)

→ More replies (5)

5

u/DigThatData Llama 7B 20d ago

chonky boi

4

u/Ok-Bill3318 20d ago

Will this run on my 1060?

12

u/80kman 20d ago

What is T is 2.4T? Does it mean Tiny? /s

3

u/No_Lingonberry1201 20d ago

Oof, that's a big boy!

4

u/1WildPanda 20d ago

Excited ---> Depressed @ 5060Ti 16G 😵

7

u/RickyRickC137 20d ago edited 20d ago

After loading this model, I’m not sure what I’m going to do with all the VRAM I have left.

6

u/JahJedi 20d ago

Use left vram for deepseek v4 pro so it will not be lonly lol.

→ More replies (2)

3

u/writeitredd 20d ago

Nothing that my dear GTX 1650 Ti cant handle.

3

u/Hefty_Wolverine_553 20d ago

Reasoning Content: Set the maximum output length to 262,144 tokens.
Final Response: Set the maximum output length to 131,072 tokens.

Using up the entirety of Qwen3.6 27B's context window for reasoning alone...

3

u/exaknight21 20d ago

I need me a 1 bit awq, that is then 1 bit’d again. Or a 4 bit that is 1 bit’d.

→ More replies (2)

3

u/fooo12gh 20d ago

If the model parameters increase at such a rate, I doubt we'll see RAM prices dropping down anytime soon.

3

u/Dance-Till-Night1 20d ago

When 3.8 30b a3b

8

u/JsThiago5 20d ago

Why hype this when 99.999% of people cannot run it? I was really expecting 27b today :(

13

u/banana_slurp_jug 20d ago

Firstly, 27b is in less than 48 hours. Secondly, even if you can't run an open-weights model on your own computer, the prices to pay for inference with it is cheaper since multiple providers will compete for value.

3

u/GregAbeI 20d ago

I think way fewer than 1 in 100,000 people can run this, which is what 99.999% means.

99.999999% of people can't run this.

→ More replies (7)

4

u/amy-schumer-tampon 20d ago

I have the feeling that a 2.4T model isn't something many people can run locally.

3

u/VR-Tech 20d ago

same for kimi, but there are people running them

2

u/SnooPaintings8639 20d ago

Big if true. Especially at full precision.

2

u/ocean_protocol 20d ago

this is wild, 95B active out of 2.4T total is a pretty aggressive sparsity ratio. anyone know what the expert routing setup looks like on this one? And how it compares to deepseek's approach

2

u/lacerating_aura 20d ago

Welp, no vision at max.

2

u/madjesta 20d ago

Could this even be run offline in any practical way just to distill? Maybe per layer batching? JFC this is huge.

2

u/Then_Blueberry7290 20d ago

text only? No vision capabilities? Or just the unsloth version doesn't contain? AFAIK qwen 3.8 27b will be vision enabled model, or not?
Anyway q1 is 400GB size...

2

u/Legitimate-Dog5690 20d ago

Amazing stuff, so glad to have this at home. Only need 5x RTX 6000s and I can run Q1.

→ More replies (2)

2

u/xXDennisXx3000 20d ago

The software is developing faster than the hardware.

2

u/xXDennisXx3000 20d ago

When Kimi K3 was a tank, Qwen 3.8 Max is a spaceship....

→ More replies (1)

2

u/Beltalowdamon 20d ago

Guessing it will be a while until you can run a distilled version of this on 8gb vram and 32gb ram!

2

u/pulse77 20d ago

Weights were uploaded 4 days ago...

2

u/a_beautiful_rhind 20d ago

This is what your "MOE" api models look like. Opus, gemini pro, etc.