r/LocalLLaMA Jul 27 '26

News Kimi K3 weights now released.

Post image

Kimi K3 weights are finally released!

3.3k Upvotes

661 comments sorted by

675

u/Simple_Split5074 Jul 27 '26

OMFG its 104B activated params

344

u/FoxiPanda Jul 27 '26

This was my general reaction too lol. 2.8T-A104B is insane lol... I'm going to admit defeat on this one and say I can't run it. You need an 8-way B300 or MI350X or a Rubin NVL8 or a cluster thereof to actually run this. What a beast.

150

u/Thomas-Lore Jul 27 '26

I was going to make a joke that I can fit one expert in my 64GB of RAM. But nope, not even that. :)

71

u/throw123awaie Jul 27 '26

They released it in MXFP4 so with around 55GB RAM you could!

→ More replies (8)
→ More replies (3)

116

u/VampiroMedicado Jul 27 '26

550k USD to run this lol

86

u/[deleted] Jul 27 '26

[removed] — view removed comment

55

u/VeterinarianOne1349 Jul 27 '26

Doesn't really work that well. This 550k setup wouldn't allow a lot of developers to work in parallel, while sitting idle during non-work hours. Makes much more sense to pay a 3rd-party to host and pay per token.

40

u/crusaderky Jul 27 '26

waiting for large corpos to rent their hardware on vast.ai during nighttime, only to find the next morning that someone ran a container jailbreak and ran wild on their private networks

→ More replies (1)
→ More replies (4)

19

u/SignificanceFlat1460 Jul 27 '26

Question: how would this scale though? Like how many units would it be required for.. let's say a group of 100 software engineers who needs it quite frequently?

→ More replies (1)

7

u/Spectrum1523 Jul 27 '26

The advantage is not running it yourself, it's that a marketplace of services will come up to run it at the lowest possible cost, and the model can't be taken offline by a single arbitrary decision

→ More replies (5)

8

u/Galdoren Jul 27 '26

The company I'm working is paying slightly over $250k per week to API costs. so yeah, 550k investment to cut the cost of the inference can be beneficial for them...

7

u/baba_bholanath Jul 28 '26

We do around 1 mil per month for OpenAI only, dont have number for Anthropic but it would be 2-3x of that given all of our use cases are around coding and agents, no wonder Anthropic is shitting their pants on open weights models, I work in Enterprise Agentic team and we have recently started fine tuning > 100 B models for specific use cases of our clients, open weights hurts Anthropic more due to enterprise customers

→ More replies (1)
→ More replies (4)
→ More replies (5)

29

u/OverclockingUnicorn Jul 27 '26

More like 2 8x nodes of B200/B300 if you actually want some context. Think it's just under 1.5TB w/o context

19

u/TheDailySpank Jul 27 '26

How many 4060-16GBs is that?

40

u/OverclockingUnicorn Jul 27 '26

200+ lol

21

u/positivitittie Jul 27 '26

Oh good. I got 3090s.

7

u/Vast_Mousse_310 Jul 27 '26

One, with a little bit of GPU offload.

→ More replies (2)
→ More replies (11)

88

u/Iwaku_Real Jul 27 '26

Holy shit that's got to be a new record too. That's like activating a new dense model 50% larger than Llama 70B for every single token. I thought it would be sparser tbh

59

u/my_name_isnt_clever Jul 27 '26

There is the tinest glimmer of hope that I could run this behemoth on my Strix Halo 128GB with the inactive weights on SSD. 1 token a minute here I come!

12

u/burritoresearch Jul 27 '26

More like 1 token every 45 minutes.

7

u/droptableadventures Jul 27 '26

104B active, weights natively in MXFP4 = gives us ~50GB of model to be read per token generation.

Let's say ~8GB/sec for the SSD. So that'd be about 1 token every 6 seconds (0.16 T/s), or 10 tokens/minute.

4

u/Head_Boysenberry5233 Jul 27 '26

i feel like 1/min would be pretty accurate based on the colibri glm 5.2 q4 implementation, about 6x slower?

Even 1 tok/min on 32gb cpu ram would be incredible and extremely useful

29

u/TechExpert2910 Jul 27 '26

let us know the perf if you try lol

8

u/RuiRdA Jul 27 '26

K3 Colibri engine lets gooo!!!

→ More replies (1)
→ More replies (6)

14

u/killerstreak976 Jul 27 '26

Holy crap, ACTIVATED params is insane

→ More replies (11)

125

u/DataGOGO Jul 27 '26

So it will run on 8 B300's in 4 bit. Pretty impressive.

63

u/Iwaku_Real Jul 27 '26

Yeah so if it were a Steam game, HGX B300 would be the recommended requirements. That's $500K of hardware (and yes it IS "local" because anyone with that amount of money could buy one to run at home)

32

u/PrinceOfLeon Jul 27 '26

Certified for Steam Deck!

6

u/No-Dot-6573 Jul 27 '26

Someone on r/SteamDeck will say it runs flawlessly.

→ More replies (1)

30

u/DataGOGO Jul 27 '26 edited Jul 28 '26

it is 1.54TB of just weights in 4 bit, you are looking at about 2TB of vram in operation,

That is roughly:

  • 86 RTX 4090 (no 4 bit accel)
  • 64 RTX 5090 ~ $450k (8 servers x 8 cards)
  • 22 RTX Pro 6000 Blackwell ~ $350k (3 severs, max 8 GPU per)
  • 16 H200 NVL (141GB) (no 4 bit accel) ~$550k (2 servers, max 8 GPU per)
  • 16 DGX Sparks ~65k (if you could get a cluster of 16 running with just 200Gb/s nics, not sure; but it would be SLOW AF)
  • 8 HGX B300's. ~$550k (1 server, 8 GPU)

Obviously not including the switches and cabling for the clusters.

24

u/wren6991 Jul 27 '26

you are looking at about 2GB of vram in operation

Perfect, this'll run great on my laptop's 4050

7

u/Iwaku_Real Jul 27 '26

You could also do HGX B200 with CPU offload since they have a shit ton of RAM too, and it would still be really fast.

→ More replies (1)

7

u/snmnky9490 Jul 27 '26

Do you mean terabytes?

→ More replies (12)
→ More replies (2)
→ More replies (2)

830

u/tonight_we_make_soap Jul 27 '26

How do I download ram in hugging face?

288

u/RevolutionaryGold325 Jul 27 '26

hf download ram

91

u/WifeyCallsMeLazy Jul 27 '26

Shhh....there is hidden flag -v for vram. I'm entrusting you to keep this secret.

26

u/ReadyAimTranspire Jul 27 '26

You wouldn't download a RAM would you?

Yes. Yes I would.

6

u/goodb1b13 Jul 27 '26

Baaaaaah!

→ More replies (1)
→ More replies (3)
→ More replies (3)

86

u/secrook Jul 27 '26

OpenAI’s latest model will hack it for you

17

u/-gh0stRush- Jul 27 '26

Thinking

Hmm, the user wants me to obtain compute resources for them. SpaceX has GPUs at their facilities at their Colossus datacenter, let me try to access those. Guessing login credentials elonmusk/420blazeitDarkMAGA...

→ More replies (1)
→ More replies (2)

34

u/Thalesian Jul 27 '26

Step 1: sign up for Google Drive
Step 2: set up a ~5 Tb instance. Will cost you
Step 3: set that cloud as your swap disk
Step 4: point kimi to use that
Step 5: enjoy your newfound independence

50

u/AmbericWizard Jul 27 '26

one token per day

19

u/Force88 Jul 27 '26

Hey, if he has good internet connection, maybe he can achieve 2-3t/d

→ More replies (1)
→ More replies (1)
→ More replies (3)

8

u/Wide-Opportunity-582 Jul 27 '26

you can download it from here

ram.exe

→ More replies (1)
→ More replies (9)

295

u/InnerLightnesses Jul 27 '26

They actually did it. Now we hope it doesn't get banned.

110

u/dennisler Jul 27 '26

that will only happen in one country i guess... while they are copying as much as possible if the technology

27

u/ChocomelP Jul 27 '26

I'm on the edge of my seat here. If the technology what?

→ More replies (5)
→ More replies (1)

30

u/itchylol742 Jul 27 '26

how would such a ban be enforced? people and small businesses even in countries that care about copyright use pirated software which is already illegal and has been for a long time, and almost never get caught

23

u/void-wanderer- Jul 27 '26

"small business", exactly. But no big corporation will risk it. And no business based on open models can be built. 

→ More replies (1)
→ More replies (3)
→ More replies (7)

260

u/de4dee Jul 27 '26

63

u/AlexanderDoak Jul 27 '26

Can I just torrent like 1% of it? You know, pitch in to show my support...

28

u/console_pleb_36935 Jul 27 '26

Yes, torrent clients will let you do that and seed a small piece.

38

u/Charl1eBr0wn Jul 27 '26

Yeah, pause it at 1%. You'd still seed depending on the client and settings (most do).

9

u/Clairvoidance Jul 27 '26

You can even choose which files you download, torrenting is a very useful format

17

u/pier4r Jul 27 '26

this, we need a p2p backup of hf

8

u/Ginden Jul 27 '26

We generally need content-adressable storage with widespread support.

There is lots of stuff that would explicitly benefit from p2p sharing, but owners have no foolproof way to provide a proper torrent, and very few people would use it.

Metalink was an interesting attempt at this, but never got popularity and tooling.

→ More replies (2)

7

u/AdDizzy8160 Jul 27 '26

... fast, s*xy, and incredibly important!!

→ More replies (9)

145

u/BlueSwordM llama.cpp Jul 27 '26

OK, I now see why Kimi K3 is so strong: it's the first open weights model in a long time to have >72B active weights

Kimi K3 is a 2.8T-A104B MoE model, damn.

14

u/stddealer Jul 27 '26

There have been some dense models with over 100B params though.

25

u/annodomini Jul 27 '26

Mistral Medium is 128B dense. And yet it performs at around the level of Gemma 4 31B. Not exactly a great tradeoff. I ran it once at one or two tokens per second and then deleted it.

7

u/BlueSwordM llama.cpp Jul 27 '26

Yes, but never an MoE from an open weights lab. I've been speculating that one of the reasons the closed weights lab have been increasing in performance more rapidly has to do with better training, but most importantly, much larger active parameters and better harnesses.

10

u/[deleted] Jul 27 '26

[deleted]

→ More replies (1)
→ More replies (2)

414

u/Blues520 Jul 27 '26

My 3090 is ready

164

u/Enfiznar Jul 27 '26

So is my 1080

106

u/SnooPaintings8639 Jul 27 '26

And my Celeron

80

u/false79 Jul 27 '26

And my abacus 

60

u/Maybe-monad Jul 27 '26

And my axe

20

u/fauxpasiii Jul 27 '26

How much VRAM your axe has?

21

u/Maybe-monad Jul 27 '26

It increases with the number of chips you smash into pieces. Right now id 6969GB.

3

u/Infinite100p Jul 28 '26

Ah, the horizontal sharding.

→ More replies (1)

5

u/Protheu5 Jul 28 '26

The axe forgets, but the tree remembers. And axe can hit multiple trees, so theoretically unlimited VRAM thanks to the axe.

11

u/NTDLS Jul 27 '26

You have an abacus? I bet you bought it before the bubble caused the prices to skyrocket. 😭

→ More replies (1)
→ More replies (4)
→ More replies (1)

51

u/TheTerrasque Jul 27 '26

My C64 is all fired up!

54

u/Novel_Friendship913 Jul 27 '26

My ESP32 already plugged into USB!!!

22

u/shankey_1906 Jul 27 '26

So is my Raspberry Pi!

21

u/nick_ziv Jul 27 '26

My copper wire is in the outlet!

12

u/MeretrixDominum Jul 27 '26

My copper wire is in my potato!

9

u/BatOk7254 Jul 27 '26

My potato is in my kitchen!

11

u/USBhost Jul 27 '26

My potato is in my garden.

→ More replies (1)

5

u/p3r3lin Jul 27 '26

Joining the rbpi army! 🫡

13

u/debackerl Jul 27 '26

My TI-85 (Z80) is hot!

→ More replies (1)

4

u/masterlafontaine Jul 27 '26

Don't forget to set an aggressive zram profile!

13

u/grav3d1gger Jul 27 '26

Mine too! I bought a 90 minute cassette tape and it’s rewound ready to go!

7

u/bitflip Jul 27 '26

You need at least a 1541 and two floppy disks for a model this size.

5

u/Holiday-Pack3385 Jul 27 '26

Heh, remember the tape drive on those? Mine had one. I can't even imagine how long it would take to load up even a 9B off that... o.O

6

u/overand Jul 27 '26

The standard ROM routine for C64 datasettes was 300 baud. We'll be generous and assume you've got a Turbo loader that'll do 3600 baud. (We'll also assume you're using a ~3GB quant of that 9B)

Load time (or, really, transfer time) would be about 77 days. Or, maybe more importantly, it would be about 930 cassettes, if my math was right. (Or, actually, other math suggests it would be about 2000 cassettes, so, IDK! Either way, it's a lot.)

→ More replies (1)

8

u/Michaeli_Starky Jul 27 '26

My analog watch is ready

6

u/screenslaver5963 Jul 27 '26

My 9070 XT is burning… wait fuck!

6

u/BatOk7254 Jul 27 '26

My 2xP40 are smoking in anticipation!

4

u/ComplexType568 Jul 27 '26

I think they're smoking for a different reason...

→ More replies (7)

90

u/Comfortable-Rock-498 Jul 27 '26

This is big for companies that want to host on-prem too. Back of the envelope calculation (could be off, correct me if I am)

If you are a large enough company that spends million+ on inference a month, it makes sense to buy a GB300 rack ($6M on top range from what I could find) which has 20.7 TB. Since the model is mixed trained (MXFP4), you would need less than 10% of the rack's memory to serve the full model. Aggregate HBM bandwidth: 576 TB/s. You can run over 6000 parallel agentic workflows (each with ~100k context on average) at ~30 tok/s.

Assuming the annual amortization+electricity at $1.5M/year and about 50% average annual utilization, you get less than 60 cents (USD) per million output token, for a frontier model with plenty of capacity to share, all your data never leaving premises and well over an order of magnitude cheaper!

38

u/autisticit Jul 27 '26

So what you mean in reality is that if 6000 of us each give $1000, we can each run Kimi 3 at 30 tok/s for cheaper than anything else?

18

u/Comfortable-Rock-498 Jul 27 '26

Tbh I have been thinking about this for months now. About how workable the co-op model is. You would need some party to do admin and maintenance. I think this might be a good business idea too - being that party who facilitates such private inference racks (billing etc is trivial)

8

u/ManIkWeet Jul 27 '26

It's not private when it's a party running it lol

9

u/Comfortable-Rock-498 Jul 27 '26

Well you do need someone for server maintenance, bills, other co-ordination etc - whatever you label it

→ More replies (8)

10

u/IgnoranceIndicatorMa Jul 27 '26

i like the way you think

4

u/pinkwar Jul 27 '26

Where do I sign?

→ More replies (3)

7

u/tempedbyfate llama.cpp Jul 27 '26

Not just corporations, I think there are nation states that are setting up their own private servers to run this for all their sensitive data.

4

u/Izento Jul 27 '26

Good number crunch. $0.60 per M is a pretty good deal

→ More replies (1)
→ More replies (1)

79

u/BarisSayit Jul 27 '26

100B active params? Damn.

75

u/noneabove1182 Bartowski Jul 27 '26

Sorry friends, but I don't think I'll be making this one :')

I don't even have enough STORAGE to hold this thing, nevermind the RAM haha

15

u/Ninjam5 Jul 27 '26

EVEN THE GOAT GAVE UP HAHA

→ More replies (1)
→ More replies (1)

268

u/nomorebuttsplz Jul 27 '26

first truly frontier open model than I cannot run on my 512 gb studio. Onward and upward!

25

u/Front_Eagle739 Jul 27 '26

Yup. Same. I think I can do about q1.5 on the macbook and mac studio combined. Im currently pondering the wisdom of one of these colibri like stream setups and using the 640GB I do have as a hot cache

3

u/nomorebuttsplz Jul 27 '26

Colibri would require being able to fit on ram for decent speed, no?

→ More replies (2)
→ More replies (2)

8

u/Square_Alps1349 Jul 27 '26

Man I’m so jealous rn. Mac Studio is nerfed at 96GB unified ram max, and the price is up 20%. So much for that 25% discount interns get. 😢 

→ More replies (1)

4

u/Tank_Gloomy Jul 27 '26

Can you try asking GPT 5.6 Sol, Opus 5 or GLM 5.2 to port it into Colibri? I don't even have the infra to say I tried, but it probably works.

→ More replies (3)
→ More replies (2)

33

u/MikeRoz Jul 27 '26

MXFP4 weights / MXFP8 activations (quantization-aware training)

So if 4-bit is 1.56 TB, then 2-bit would be roughly 798 GB?

13

u/habibyajam Llama 405B Jul 27 '26

So I need a 0.03-bit quant to run it fully on my GPU. Nice!

32

u/Few_Painter_5588 Jul 27 '26

Holy shit, 104B active paramaters???

→ More replies (6)

33

u/TheRealMasonMac Jul 27 '26

They have a new license:

If the Licensee or any of its affiliates operates a Model as a Service business, and the aggregate revenue of the Licensee and its affiliates exceeds 20 million US dollars (or the equivalent in other currencies) in total over any consecutive 12 months, the Licensee must enter into a separate agreement with Moonshot AI before using the Software or its derivative works for any commercial purpose.

18

u/mthmchris Jul 27 '26

My guess is this is what MOFCOM was trying to thread the needle with. I.e. trying to keep upper leadership’s commitment to open source, while avoiding, say, MSFT just grabbing it and tweaking it slightly to have an “American version” (while banning the “Chinese versions”).

It kinda sucks, but I get the logic.

7

u/TheRealMasonMac Jul 27 '26

I suspect it also means that providers won't be able to undercut them as aggressively (if at all). Not necessarily because Moonshot says so, but presumably because of any requirements they put on providers and any royalties they expect. MiniMax, for instance, limited some providers to only using B200 or B300 GPUs.

3

u/Venryx Jul 27 '26

Wait, how are there already five other providers of the model on OpenRouter then? Surely they haven't all already made individual deals with Moonshot?

6

u/TheRealMasonMac Jul 27 '26

They had six official partners for launch (you can see on their Twitter). Those five on OpenRouter are indeed partners. But it's part of their license agreement that if you are over a certain revenue you must have a contract with them.

→ More replies (1)

206

u/durden111111 Jul 27 '26

my 512mb integrated graphics is so fucking ready

73

u/SavunOski Jul 27 '26

Negative tokens/s, you're gonna suck away the tokens

6

u/nanihikaru01 Jul 27 '26

Free money hack you say?

4

u/learn_and_learn Jul 27 '26

Sounds freaky

5

u/Iwaku_Real Jul 27 '26

To find the answer to life the universe and everything?

20 days of prompt processing: 4

Probably 10 days after that: 2

6

u/m0j0m0j Jul 27 '26

Need 0 bit quantized version for that

→ More replies (2)

55

u/SnooPaintings8639 Jul 27 '26

Who's gonna be the first brave soul to measure tps when streaming from hard drive?

35

u/Front_Eagle739 Jul 27 '26

Sigh. Im going to have to try from pure curiosity.  600GB of hot cache on the macs, another TB streaming from ssd to my rtx 5090. This can only be a good idea (farewell my next few days of productivity)

8

u/Ninjam5 Jul 27 '26

Update?

6

u/Front_Eagle739 Jul 27 '26

My Internet is very slow,  results pending

→ More replies (1)
→ More replies (1)
→ More replies (6)

63

u/HulksInvinciblePants Jul 27 '26

Kimi K3 27B when?

23

u/Browserurd Jul 27 '26

If someone can distill it to Qwen 122B that would be super.

→ More replies (1)

23

u/Top-Handle-5728 Jul 27 '26

Leave vram I do not even have the disk storage to use this model. A few with storage can dare to use AirLLM for experiencing the intergalactic streaming of voyager at 160 bits a second. Even that seems pretty fast ig

20

u/vr_fanboy Jul 27 '26

a month ago we were told that a new jump in intellegence was made, 'mythos' class models were born, too dangerous for us plebs. Forward a month, we have an open source 'mythos' class model, acceleration or anthropics regular bullshit?

Btw dont understand markets, DS3 destroyed the stocks and this does...nothing. This feels more significant if more people can serve the same drugs as oai or anthropic on the cheap.

→ More replies (2)

65

u/just_a_fan123 Jul 27 '26

Can this run on a single DGX spark at 0.5B quant?

41

u/SavunOski Jul 27 '26

Of course. Wonderfully might I add

23

u/THESALTEDPEANUT Jul 27 '26

I'm new around here and I can't tell if this thread is all sarcasm or not. 

36

u/SavunOski Jul 27 '26

Sorry, yeah that was sarcastic, forgot to add /s. Many people make jokes here, so I kinda forgot

11

u/THESALTEDPEANUT Jul 27 '26

I kinda figured but it's not just you it's like every comment lol, appreciate it though. 

8

u/PomegranateGreen3698 Jul 27 '26

Ha, I think it's the nature of this release. Largest open weight model ever released, doesn't really have any possibility for "Local Usage" bc it'd cost something like $500k in GPUs.

→ More replies (1)
→ More replies (1)

59

u/THE--GRINCH Jul 27 '26

my laptop rtx 2050 is ready to throw hands

112

u/sumane12 Jul 27 '26

Even if you cant run it, download it.

95

u/DeProgrammer99 Jul 27 '26

Can't even do that...it's the same size as the total used space on my SSD.

45

u/ChampionshipIcy7602 Jul 27 '26

Can't even download it lmao

→ More replies (1)

30

u/Herr_Drosselmeyer Jul 27 '26

Why waste terabytes worth of space for something I will never use? 

70

u/seg_lol Jul 27 '26

Trade it for antibiotics in the apocalypse.

→ More replies (8)

16

u/some_user_2021 Jul 27 '26

I remember downloading huge N64 ROMs that took loads of space and no emulator could run. Running those ROMs now is trivial.

20

u/Herr_Drosselmeyer Jul 27 '26

So is downloading them.

→ More replies (1)

6

u/banana_slurp_jug Jul 27 '26

I have 1TB total, don't even have a hard drive big enough to store the whole thing in my house.

6

u/No_Conversation9561 Jul 27 '26

someone download it, compress it and upload it and then I will download it

→ More replies (7)

11

u/Mindless_Selection34 Jul 27 '26

how much does it weight

45

u/SavunOski Jul 27 '26

2.8T parameters with 104B active. The files take 1.56TB of space to download.

11

u/meca23 Jul 27 '26

Oh so it's quantized as 4 bits?

7

u/zkstx llama.cpp Jul 27 '26

Yes, there is also a tech report. QAT from SFT phase onwards

8

u/Lissanro Jul 27 '26

I wish I had two TB of RAM instead of just one. I guess I will have to wait for Q2 GGUF to run it on my workstation. Still, will be interesting to try and see how it's Q2 quant compairs against Kimi K2.7 Q4_X. 

→ More replies (5)
→ More replies (1)
→ More replies (3)

11

u/IamNotMike25 Jul 27 '26

Historic moment tbh

27

u/Forsaken-Mode-3422 Jul 27 '26

Finally, model i cant run, but its already cool, nice

20

u/SavunOski Jul 27 '26

Cloud prices will likely drop with competition, beneficial for everyone

8

u/Forsaken-Mode-3422 Jul 27 '26

i just hope qwen will publish not only 3.8 Max, but smaller models as well, so we can enjoy the local frontier ourself

→ More replies (1)

19

u/KenTitan Jul 27 '26

I'm so broke I don't even have enough hard drive space to download

15

u/MixtureOfAmateurs koboldcpp Jul 27 '26

Hugging face down? Lmao

Edit: Nevermind it's back. Might have been on my end ¯_(ツ)_/¯

13

u/SavunOski Jul 27 '26

Seems to be up for me. If you live in a big country, there is a possibility the local CDNs are having trouble keeping up. Especially in the US, where I imagine thousands are downloading the model just in case the government bans it later.

6

u/AdDizzy8160 Jul 27 '26

.torrent|magnet link … as soon as possible!

→ More replies (1)

6

u/Morphon Jul 27 '26

We'll probably see some new, faster providers pop up. With any luck, this will turn out to be a solid distillation parent. I'd be interested to see if we get some good downstream models that can run on consumer hardware.

15

u/PerfectOlive1324 Jul 27 '26

Hoping for an unsloth Q0_XS quant I can run locally 🙏

6

u/Inevitable_Mistake32 Jul 27 '26

UD_IQ0_XS Pls and ty.

→ More replies (1)

14

u/AdDizzy8160 Jul 27 '26

Ok, thanx moonshot team!

12

u/Informal-Trouble2183 Jul 27 '26

Inference providers are going to have a good business

16

u/SavunOski Jul 27 '26

Infinite demand, literally

→ More replies (7)

5

u/ReasonablePossum_ Jul 27 '26

Hope this opens the model in more providers soon ,because its a dan pain to get a single response from the kimi app

→ More replies (1)

6

u/Hefty_Acanthaceae348 Jul 27 '26

I have high hopes that even if I can't run such a model, it will enable the creation of high quality datasets

5

u/aboutthednm Jul 27 '26

I'm starting to understand why they had "some compute problems" when first rolling it out on the API, 100B+ active params per token, it's beastly lmao. Good work.

6

u/Easy_Werewolf7903 Jul 27 '26

I want to see someone manually calculated a token by hand using Kimi K3 weights.

18

u/equatorbit Jul 27 '26

Gonna try to get this running on an ESP32

3

u/killerstreak976 Jul 27 '26

Kimi k3, meet my ESP32-C2

→ More replies (4)

10

u/Tedinasuit Jul 27 '26

What a beautiful day

5

u/DragonfruitIll660 Jul 27 '26 edited Jul 27 '26

Ayyy lets go. That's awesome news. Also holy 104B active, never seen a MoE with that many active, its almost Mistral 2 large sized.

4

u/patricious llama.cpp Jul 27 '26

Wake me when a provider has it with a good price per 1mil.

4

u/mikewilkinsjr Jul 27 '26

I can’t run it, I’ll never be able to run it.

I am also fighting the urge to download the weights and squirrel them away like an out of control data hoarder.

5

u/crusaderky Jul 27 '26

Jokes aside.
This on paper fits on a HGX B300 ($~0.5m, 2304 GB VRAM)...

...or on 24x Atlas 300I, $1300 each, $31,200 total, plus 3~4 XEON/threadrippers with 800G networking. So.... less than $60k in total?

Who's got the spare change to try?

→ More replies (2)

11

u/ilintar Jul 27 '26

Okay, time to get that 0.1Q_0 quant support rolling in llama.cpp!

8

u/milkipedia Jul 27 '26

I am not data hoarder enough to store these weights for kicks without the hardware to run it. Enjoy, folks

5

u/msew Jul 27 '26

I wonder how long until consumer hardware will have the requirements to run all these and future LLMs locally for like $5k-10k

→ More replies (6)

4

u/AdOne8437 Jul 27 '26

Well, in 15-20 years we will have consumer cards to run it.

→ More replies (1)

13

u/loversama Jul 27 '26

Quick before it gets banned 😂

10

u/WonderFactory Jul 27 '26

I dont think it's worth banning it now it's released but I was a bit anxious leading up to the release in case something stopped it from being released

→ More replies (1)

6

u/bad_detectiv3 Jul 27 '26

How much vram do I need to run this

7

u/fishslinger Jul 27 '26

If you have to ask, you can't afford it

12

u/SavunOski Jul 27 '26

Around 2TB :x

12

u/jijig Jul 27 '26

Hey if you already got 2 3090s you only need 81 more

9

u/Melbar666 Jul 27 '26

waiting for Kimi-K3-Q0-Abliterated-Heretic.GGUF

6

u/MotokoAGI Jul 27 '26

:-)

:-(

8

u/Mayion Jul 27 '26

Fullmetal Alchemist intermissions be like

9

u/jacek2023 llama.cpp Jul 27 '26

I am GPU poor I have only 4x3090 so I can't run it. I envy you all from r/LocalLLaMA with stronger setups.

→ More replies (10)