r/LocalLLaMA • u/SavunOski • Jul 27 '26
News Kimi K3 weights now released.
Kimi K3 weights are finally released!
125
u/DataGOGO Jul 27 '26
So it will run on 8 B300's in 4 bit. Pretty impressive.
→ More replies (2)63
u/Iwaku_Real Jul 27 '26
Yeah so if it were a Steam game, HGX B300 would be the recommended requirements. That's $500K of hardware (and yes it IS "local" because anyone with that amount of money could buy one to run at home)
32
→ More replies (2)30
u/DataGOGO Jul 27 '26 edited Jul 28 '26
it is 1.54TB of just weights in 4 bit, you are looking at about 2TB of vram in operation,
That is roughly:
- 86 RTX 4090 (no 4 bit accel)
- 64 RTX 5090 ~ $450k (8 servers x 8 cards)
- 22 RTX Pro 6000 Blackwell ~ $350k (3 severs, max 8 GPU per)
- 16 H200 NVL (141GB) (no 4 bit accel) ~$550k (2 servers, max 8 GPU per)
- 16 DGX Sparks ~65k (if you could get a cluster of 16 running with just 200Gb/s nics, not sure; but it would be SLOW AF)
- 8 HGX B300's. ~$550k (1 server, 8 GPU)
Obviously not including the switches and cabling for the clusters.
24
u/wren6991 Jul 27 '26
you are looking at about 2GB of vram in operation
Perfect, this'll run great on my laptop's 4050
7
u/Iwaku_Real Jul 27 '26
You could also do HGX B200 with CPU offload since they have a shit ton of RAM too, and it would still be really fast.
→ More replies (1)→ More replies (12)7
830
u/tonight_we_make_soap Jul 27 '26
How do I download ram in hugging face?
288
u/RevolutionaryGold325 Jul 27 '26
hf download ram
→ More replies (3)91
u/WifeyCallsMeLazy Jul 27 '26
Shhh....there is hidden flag -v for vram. I'm entrusting you to keep this secret.
→ More replies (3)26
u/ReadyAimTranspire Jul 27 '26
You wouldn't download a RAM would you?
Yes. Yes I would.
→ More replies (1)6
86
u/secrook Jul 27 '26
OpenAI’s latest model will hack it for you
→ More replies (2)17
u/-gh0stRush- Jul 27 '26
Thinking
Hmm, the user wants me to obtain compute resources for them. SpaceX has GPUs at their facilities at their Colossus datacenter, let me try to access those. Guessing login credentials elonmusk/420blazeitDarkMAGA...
→ More replies (1)26
34
u/Thalesian Jul 27 '26
Step 1: sign up for Google Drive
Step 2: set up a ~5 Tb instance. Will cost you
Step 3: set that cloud as your swap disk
Step 4: point kimi to use that
Step 5: enjoy your newfound independence→ More replies (3)50
u/AmbericWizard Jul 27 '26
one token per day
→ More replies (1)19
u/Force88 Jul 27 '26
Hey, if he has good internet connection, maybe he can achieve 2-3t/d
→ More replies (1)8
→ More replies (9)9
u/Iwaku_Real Jul 27 '26
https://huggingface.co/baa-ai/Qwen3.5-122B-A10B-RAM-60GB-MLX, it actually comes with 60GB RAM! /s
295
u/InnerLightnesses Jul 27 '26
They actually did it. Now we hope it doesn't get banned.
110
u/dennisler Jul 27 '26
that will only happen in one country i guess... while they are copying as much as possible if the technology
→ More replies (1)27
→ More replies (7)30
u/itchylol742 Jul 27 '26
how would such a ban be enforced? people and small businesses even in countries that care about copyright use pirated software which is already illegal and has been for a long time, and almost never get caught
→ More replies (3)23
u/void-wanderer- Jul 27 '26
"small business", exactly. But no big corporation will risk it. And no business based on open models can be built.
→ More replies (1)
260
u/de4dee Jul 27 '26
63
u/AlexanderDoak Jul 27 '26
Can I just torrent like 1% of it? You know, pitch in to show my support...
28
38
u/Charl1eBr0wn Jul 27 '26
Yeah, pause it at 1%. You'd still seed depending on the client and settings (most do).
9
u/Clairvoidance Jul 27 '26
You can even choose which files you download, torrenting is a very useful format
17
u/pier4r Jul 27 '26
this, we need a p2p backup of hf
8
u/Ginden Jul 27 '26
We generally need content-adressable storage with widespread support.
There is lots of stuff that would explicitly benefit from p2p sharing, but owners have no foolproof way to provide a proper torrent, and very few people would use it.
Metalink was an interesting attempt at this, but never got popularity and tooling.
→ More replies (2)→ More replies (9)7
145
u/BlueSwordM llama.cpp Jul 27 '26
OK, I now see why Kimi K3 is so strong: it's the first open weights model in a long time to have >72B active weights
Kimi K3 is a 2.8T-A104B MoE model, damn.
14
u/stddealer Jul 27 '26
There have been some dense models with over 100B params though.
25
u/annodomini Jul 27 '26
Mistral Medium is 128B dense. And yet it performs at around the level of Gemma 4 31B. Not exactly a great tradeoff. I ran it once at one or two tokens per second and then deleted it.
7
u/BlueSwordM llama.cpp Jul 27 '26
Yes, but never an MoE from an open weights lab. I've been speculating that one of the reasons the closed weights lab have been increasing in performance more rapidly has to do with better training, but most importantly, much larger active parameters and better harnesses.
→ More replies (2)10
414
u/Blues520 Jul 27 '26
My 3090 is ready
164
u/Enfiznar Jul 27 '26
So is my 1080
→ More replies (1)106
u/SnooPaintings8639 Jul 27 '26
And my Celeron
80
u/false79 Jul 27 '26
And my abacus
60
u/Maybe-monad Jul 27 '26
And my axe
20
u/fauxpasiii Jul 27 '26
How much VRAM your axe has?
21
u/Maybe-monad Jul 27 '26
It increases with the number of chips you smash into pieces. Right now id 6969GB.
→ More replies (1)3
5
u/Protheu5 Jul 28 '26
The axe forgets, but the tree remembers. And axe can hit multiple trees, so theoretically unlimited VRAM thanks to the axe.
→ More replies (4)11
u/NTDLS Jul 27 '26
You have an abacus? I bet you bought it before the bubble caused the prices to skyrocket. 😭
→ More replies (1)51
u/TheTerrasque Jul 27 '26
My C64 is all fired up!
54
u/Novel_Friendship913 Jul 27 '26
My ESP32 already plugged into USB!!!
22
u/shankey_1906 Jul 27 '26
So is my Raspberry Pi!
21
u/nick_ziv Jul 27 '26
My copper wire is in the outlet!
12
u/MeretrixDominum Jul 27 '26
My copper wire is in my potato!
9
5
13
4
13
u/grav3d1gger Jul 27 '26
Mine too! I bought a 90 minute cassette tape and it’s rewound ready to go!
7
→ More replies (1)5
u/Holiday-Pack3385 Jul 27 '26
Heh, remember the tape drive on those? Mine had one. I can't even imagine how long it would take to load up even a 9B off that... o.O
6
u/overand Jul 27 '26
The standard ROM routine for C64 datasettes was 300 baud. We'll be generous and assume you've got a Turbo loader that'll do 3600 baud. (We'll also assume you're using a ~3GB quant of that 9B)
Load time (or, really, transfer time) would be about 77 days. Or, maybe more importantly, it would be about 930 cassettes, if my math was right. (Or, actually, other math suggests it would be about 2000 cassettes, so, IDK! Either way, it's a lot.)
8
6
→ More replies (7)6
90
u/Comfortable-Rock-498 Jul 27 '26
This is big for companies that want to host on-prem too. Back of the envelope calculation (could be off, correct me if I am)
If you are a large enough company that spends million+ on inference a month, it makes sense to buy a GB300 rack ($6M on top range from what I could find) which has 20.7 TB. Since the model is mixed trained (MXFP4), you would need less than 10% of the rack's memory to serve the full model. Aggregate HBM bandwidth: 576 TB/s. You can run over 6000 parallel agentic workflows (each with ~100k context on average) at ~30 tok/s.
Assuming the annual amortization+electricity at $1.5M/year and about 50% average annual utilization, you get less than 60 cents (USD) per million output token, for a frontier model with plenty of capacity to share, all your data never leaving premises and well over an order of magnitude cheaper!
38
u/autisticit Jul 27 '26
So what you mean in reality is that if 6000 of us each give $1000, we can each run Kimi 3 at 30 tok/s for cheaper than anything else?
18
u/Comfortable-Rock-498 Jul 27 '26
Tbh I have been thinking about this for months now. About how workable the co-op model is. You would need some party to do admin and maintenance. I think this might be a good business idea too - being that party who facilitates such private inference racks (billing etc is trivial)
→ More replies (8)8
u/ManIkWeet Jul 27 '26
It's not private when it's a party running it lol
9
u/Comfortable-Rock-498 Jul 27 '26
Well you do need someone for server maintenance, bills, other co-ordination etc - whatever you label it
10
→ More replies (3)4
7
u/tempedbyfate llama.cpp Jul 27 '26
Not just corporations, I think there are nation states that are setting up their own private servers to run this for all their sensitive data.
→ More replies (1)4
79
75
u/noneabove1182 Bartowski Jul 27 '26
Sorry friends, but I don't think I'll be making this one :')
I don't even have enough STORAGE to hold this thing, nevermind the RAM haha
→ More replies (1)15
268
u/nomorebuttsplz Jul 27 '26
first truly frontier open model than I cannot run on my 512 gb studio. Onward and upward!
25
u/Front_Eagle739 Jul 27 '26
Yup. Same. I think I can do about q1.5 on the macbook and mac studio combined. Im currently pondering the wisdom of one of these colibri like stream setups and using the 640GB I do have as a hot cache
→ More replies (2)3
u/nomorebuttsplz Jul 27 '26
Colibri would require being able to fit on ram for decent speed, no?
→ More replies (2)8
u/Square_Alps1349 Jul 27 '26
Man I’m so jealous rn. Mac Studio is nerfed at 96GB unified ram max, and the price is up 20%. So much for that 25% discount interns get. 😢
→ More replies (1)→ More replies (2)4
u/Tank_Gloomy Jul 27 '26
Can you try asking GPT 5.6 Sol, Opus 5 or GLM 5.2 to port it into Colibri? I don't even have the infra to say I tried, but it probably works.
→ More replies (3)
33
u/MikeRoz Jul 27 '26
MXFP4 weights / MXFP8 activations (quantization-aware training)
So if 4-bit is 1.56 TB, then 2-bit would be roughly 798 GB?
13
32
33
u/TheRealMasonMac Jul 27 '26
They have a new license:
If the Licensee or any of its affiliates operates a Model as a Service business, and the aggregate revenue of the Licensee and its affiliates exceeds 20 million US dollars (or the equivalent in other currencies) in total over any consecutive 12 months, the Licensee must enter into a separate agreement with Moonshot AI before using the Software or its derivative works for any commercial purpose.
18
u/mthmchris Jul 27 '26
My guess is this is what MOFCOM was trying to thread the needle with. I.e. trying to keep upper leadership’s commitment to open source, while avoiding, say, MSFT just grabbing it and tweaking it slightly to have an “American version” (while banning the “Chinese versions”).
It kinda sucks, but I get the logic.
7
u/TheRealMasonMac Jul 27 '26
I suspect it also means that providers won't be able to undercut them as aggressively (if at all). Not necessarily because Moonshot says so, but presumably because of any requirements they put on providers and any royalties they expect. MiniMax, for instance, limited some providers to only using B200 or B300 GPUs.
→ More replies (1)3
u/Venryx Jul 27 '26
Wait, how are there already five other providers of the model on OpenRouter then? Surely they haven't all already made individual deals with Moonshot?
6
u/TheRealMasonMac Jul 27 '26
They had six official partners for launch (you can see on their Twitter). Those five on OpenRouter are indeed partners. But it's part of their license agreement that if you are over a certain revenue you must have a contract with them.
206
u/durden111111 Jul 27 '26
my 512mb integrated graphics is so fucking ready
73
5
u/Iwaku_Real Jul 27 '26
To find the answer to life the universe and everything?
20 days of prompt processing:
4Probably 10 days after that:
2→ More replies (2)6
55
u/SnooPaintings8639 Jul 27 '26
Who's gonna be the first brave soul to measure tps when streaming from hard drive?
→ More replies (6)35
u/Front_Eagle739 Jul 27 '26
Sigh. Im going to have to try from pure curiosity. 600GB of hot cache on the macs, another TB streaming from ssd to my rtx 5090. This can only be a good idea (farewell my next few days of productivity)
8
63
23
u/Top-Handle-5728 Jul 27 '26
Leave vram I do not even have the disk storage to use this model. A few with storage can dare to use AirLLM for experiencing the intergalactic streaming of voyager at 160 bits a second. Even that seems pretty fast ig
20
u/vr_fanboy Jul 27 '26
a month ago we were told that a new jump in intellegence was made, 'mythos' class models were born, too dangerous for us plebs. Forward a month, we have an open source 'mythos' class model, acceleration or anthropics regular bullshit?
Btw dont understand markets, DS3 destroyed the stocks and this does...nothing. This feels more significant if more people can serve the same drugs as oai or anthropic on the cheap.
→ More replies (2)
65
u/just_a_fan123 Jul 27 '26
Can this run on a single DGX spark at 0.5B quant?
→ More replies (1)41
u/SavunOski Jul 27 '26
Of course. Wonderfully might I add
23
u/THESALTEDPEANUT Jul 27 '26
I'm new around here and I can't tell if this thread is all sarcasm or not.
36
u/SavunOski Jul 27 '26
Sorry, yeah that was sarcastic, forgot to add /s. Many people make jokes here, so I kinda forgot
11
u/THESALTEDPEANUT Jul 27 '26
I kinda figured but it's not just you it's like every comment lol, appreciate it though.
8
u/PomegranateGreen3698 Jul 27 '26
Ha, I think it's the nature of this release. Largest open weight model ever released, doesn't really have any possibility for "Local Usage" bc it'd cost something like $500k in GPUs.
→ More replies (1)
59
112
u/sumane12 Jul 27 '26
Even if you cant run it, download it.
95
u/DeProgrammer99 Jul 27 '26
Can't even do that...it's the same size as the total used space on my SSD.
45
30
u/Herr_Drosselmeyer Jul 27 '26
Why waste terabytes worth of space for something I will never use?
70
16
u/some_user_2021 Jul 27 '26
I remember downloading huge N64 ROMs that took loads of space and no emulator could run. Running those ROMs now is trivial.
→ More replies (1)20
6
u/banana_slurp_jug Jul 27 '26
I have 1TB total, don't even have a hard drive big enough to store the whole thing in my house.
→ More replies (7)6
u/No_Conversation9561 Jul 27 '26
someone download it, compress it and upload it and then I will download it
11
u/Mindless_Selection34 Jul 27 '26
how much does it weight
→ More replies (3)45
u/SavunOski Jul 27 '26
2.8T parameters with 104B active. The files take 1.56TB of space to download.
11
→ More replies (1)8
u/Lissanro Jul 27 '26
I wish I had two TB of RAM instead of just one. I guess I will have to wait for Q2 GGUF to run it on my workstation. Still, will be interesting to try and see how it's Q2 quant compairs against Kimi K2.7 Q4_X.
→ More replies (5)
11
27
u/Forsaken-Mode-3422 Jul 27 '26
Finally, model i cant run, but its already cool, nice
→ More replies (1)20
u/SavunOski Jul 27 '26
Cloud prices will likely drop with competition, beneficial for everyone
8
u/Forsaken-Mode-3422 Jul 27 '26
i just hope qwen will publish not only 3.8 Max, but smaller models as well, so we can enjoy the local frontier ourself
19
15
u/MixtureOfAmateurs koboldcpp Jul 27 '26
Hugging face down? Lmao
Edit: Nevermind it's back. Might have been on my end ¯_(ツ)_/¯
13
u/SavunOski Jul 27 '26
Seems to be up for me. If you live in a big country, there is a possibility the local CDNs are having trouble keeping up. Especially in the US, where I imagine thousands are downloading the model just in case the government bans it later.
6
6
u/Morphon Jul 27 '26
We'll probably see some new, faster providers pop up. With any luck, this will turn out to be a solid distillation parent. I'd be interested to see if we get some good downstream models that can run on consumer hardware.
15
u/PerfectOlive1324 Jul 27 '26
Hoping for an unsloth Q0_XS quant I can run locally 🙏
→ More replies (1)6
14
12
u/Informal-Trouble2183 Jul 27 '26
Inference providers are going to have a good business
→ More replies (7)16
5
u/ReasonablePossum_ Jul 27 '26
Hope this opens the model in more providers soon ,because its a dan pain to get a single response from the kimi app
→ More replies (1)
6
u/Hefty_Acanthaceae348 Jul 27 '26
I have high hopes that even if I can't run such a model, it will enable the creation of high quality datasets
5
u/aboutthednm Jul 27 '26
I'm starting to understand why they had "some compute problems" when first rolling it out on the API, 100B+ active params per token, it's beastly lmao. Good work.
6
u/Easy_Werewolf7903 Jul 27 '26
I want to see someone manually calculated a token by hand using Kimi K3 weights.
18
10
5
u/DragonfruitIll660 Jul 27 '26 edited Jul 27 '26
Ayyy lets go. That's awesome news. Also holy 104B active, never seen a MoE with that many active, its almost Mistral 2 large sized.
4
4
u/mikewilkinsjr Jul 27 '26
I can’t run it, I’ll never be able to run it.
I am also fighting the urge to download the weights and squirrel them away like an out of control data hoarder.
5
u/crusaderky Jul 27 '26
Jokes aside.
This on paper fits on a HGX B300 ($~0.5m, 2304 GB VRAM)...
...or on 24x Atlas 300I, $1300 each, $31,200 total, plus 3~4 XEON/threadrippers with 800G networking. So.... less than $60k in total?
Who's got the spare change to try?
→ More replies (2)
11
8
u/milkipedia Jul 27 '26
I am not data hoarder enough to store these weights for kicks without the hardware to run it. Enjoy, folks
5
u/msew Jul 27 '26
I wonder how long until consumer hardware will have the requirements to run all these and future LLMs locally for like $5k-10k
→ More replies (6)
4
u/AdOne8437 Jul 27 '26
Well, in 15-20 years we will have consumer cards to run it.
→ More replies (1)
13
u/loversama Jul 27 '26
Quick before it gets banned 😂
10
u/WonderFactory Jul 27 '26
I dont think it's worth banning it now it's released but I was a bit anxious leading up to the release in case something stopped it from being released
→ More replies (1)
6
u/bad_detectiv3 Jul 27 '26
How much vram do I need to run this
44
7
12
9
6
9
u/jacek2023 llama.cpp Jul 27 '26
I am GPU poor I have only 4x3090 so I can't run it. I envy you all from r/LocalLLaMA with stronger setups.
→ More replies (10)



675
u/Simple_Split5074 Jul 27 '26
OMFG its 104B activated params