r/codex 19h ago

Commentary Token prices are soon going down

I know a lot of people come to this sub just to complain about hitting their limits (in a way I honestly don't understand, unless they're letting agents run 24/7 completely unsupervised).

The thing is, these individual plans are heavily subsidized by corporate API plans, the latter are incredibly expensive compared to the token usage a single user can rack up on an individual plan.

Companies are already strictly limiting token consumption based on the size of the bill they receive at the end of the month.

Recently, open-weight models have been released that rival the top-tier models from OpenAI and Anthropic in terms of task performance and reasoning. Hosting these models locally, with token/second speeds acceptable for up to 100 employees using them simultaneously, is actually affordable when compared to long-term API costs. A server capable of running the 2.8-trillion-parameter Kimi K3 model costs less than $2 million, and if resources are managed wisely, it is possible to acquire a machine with these capabilities for under $1 million.

16x NVidia H200 to 2.8TB VRAM, current prices between $640k and $840k only for the GPUs.

Kimi K3 2.8 trillion parameters to ~1.56TB for the weights

Double that to KV Cache 100 users at the same time.

Obviously, the cost rises significantly due to the need to hire staff to manage these machines, along with electricity bills and everything else. Even so, it is advantageous to do this for the long term, in addition to the fine-tuning capability, which greatly increases productivity

Major companies will soon start migrating to this, which will force the big AI ones to develop methods to lower the cost of processing models, thereby reducing hosting costs for them, and compel them to offer competitive pricing.

I expect this to trickle down to consumers by the end of 2028.

All hail open-weight models!

0 Upvotes

30 comments sorted by

4

u/Tetriandoch 18h ago

Following you logic, big companies won't subsidize our little individual plans anymore, because they move to in-house solutions. If anything, this would mean that OpenAI has to increase prices or decrease usage limits for us. They are always working on more efficient models regardless, because it increases their profit margin. They already have big incentive to make models more efficient and/or attract more customers because they are not yet profitable. If open-source models are taking away customers, they would actually have to increase prices. People would pay for it because they offer the most capable models.

I appreciate the discussion but I am not sure your logic is sound.

-5

u/Fit-Gas-5760 18h ago edited 18h ago

If open-source models are taking away customers, they would actually have to increase prices. People would pay for it because they offer the most capable models.

This is a suicide sentence in a free-market economy, they will do this only if they want to stack a lot of money right before selling the company.

And as I said, Kimi K3 is currently on par with Sol and Opus 5.

1

u/SourSovereign 18h ago

Well then it's a lose-lose situation.

Either lose more money by stoping subsidiaries and losing subscribers or continue with it and still lose money because of it.

9

u/SlimyResearcher 19h ago

#slop

-2

u/Fit-Gas-5760 19h ago

? If you disagree you can join the discussion... but this wont add anything to it

5

u/RealSlyck 19h ago

Token prices are subsidized right now by tons of VC money, son. APIs, subs, and ads are just footnotes for valuation.

Also, root comment holds true. Original piece is War and Peace, your response isn’t even a haiku. Slop is as slop does.

1

u/Fit-Gas-5760 18h ago

It has been some time since AI companies started turning a profit and stopped relying solely on VC.

Also, root comment holds true. Original piece is War and Peace, your response isn’t even a haiku. Slop is as slop does.

Wtf?

-1

u/RealSlyck 17h ago

What? You might want to inform Wall Street and literally all humans about this, bot…this would be news for everyone if true. But we know it’s not, again, slop is as slop does.

1

u/Fit-Gas-5760 17h ago

1

u/RealSlyck 17h ago

Yeah notice your words champ. Profit != revenue != valuation. Bots don’t pay attention to detail, we know.

1

u/Fit-Gas-5760 17h ago

Nah, youre def high af. If you just had read the title of the article... but I think humans dont know how to read anymore.

2

u/RealSlyck 17h ago

Might be, but the post and thesis is slop. Hang in there kid.

-1

u/SlimyResearcher 17h ago

It's not slop, it's AI!

2

u/Factor013 18h ago

The thing is, these individual plans are heavily subsidized by corporate API plans, the latter are incredibly expensive compared to the token usage a single user can rack up on an individual plan.

Why do people keep spreading this myth? I honestly don't understand! Subs are not subsidized! What they charge per token via API pricing has nothing to do with how much it actually costs them per token.

I already explained this in another topic:

They charge high API prices because of the value they believe their service provides. In other words... when you sell a product that replaces something more costly (for example, human labor/developers etc) then you can easily charge 50% of whatever net gain your customer ends up with for using your product as that still makes it a good deal for that customer.

So that is the primary factor what decides what they can charge for their API pricing... not the actual costs.

So TDLR: When you sell something so unique and desirable because it saves your customer money in the long run... you can basically charge whatever you want for it. And that is exactly what they are doing. True costs is hardly a factor in any of this, it will be marginal.

Subscriptions will probably be more in line with the actual operational costs to provide that service... But still profitable for them due to these subs being a free source of R&D (training data).

1

u/Bladder-Splatter 18h ago

Given the speed we've moved in just a year I'd say predicting where we are in 2028 is wild, I legitimately cannot confidently say what we won't be able to do by then given this exponential growth, including costs.

I also snorted at ONLY $2 million but that's because I come from an era that kind of money was a mainframe or a fucking military bunker if you were talking buyable tech parts.

2

u/Fit-Gas-5760 18h ago edited 18h ago

"Only $2 million" in the perspective of expenses and revenue of a company with 500+ employees nowadays, $2 million is close to "only" for this scenario.

"Agentic coding" for 200 developers can easily exceed $1 million per year with OpenAI plans

1

u/dvduval 18h ago

The problem with this idea is these big corporations are not investing massively and not only the AI model, but the infrastructure is supported and a piece of hardware is not enough to be a game changer here.

If anything, they’re creating even tighter relationships with the big AI companies right now, and having dedicated teams working together.

And that’s why a lot of the AI companies are also closely aligned or even in control of the actual hardware and data centers. They have purchased way ahead of everybody else for things like ram. The big frontier models thought way ahead on all this stuff.

1

u/fibonac1123 18h ago edited 18h ago

By this logic, AWS, Azure and friends would not exist anymore, yet they print money. Some companies wil get their own LLM servers, doubt it will be even 10% a few years ahead. Convenience, easier accounting, etc. matters far more for big companies, and small companies can't affort the really powerful servers even if they had the people knowing how to set it up.

1

u/[deleted] 14h ago

[deleted]

1

u/Fit-Gas-5760 10h ago

The point is def not giving info to shareholders lol

1

u/OtherwiseAlbatross14 12h ago

Lol 2+ years in this industry is anything but "soon"

1

u/AkindaGood_programer 8h ago

The biggest problem with this post is that you're assuming that Frontier labs (OpenAI, Anthropic, etc) are releasing their best models to the public. Chinese companies almost instantly release anything they make, but frontier labs often keep models for months, fine-tuning them, ensuring safety, etc.

Mythos came out in super early April; that is almost five full months. Do you really think that within that time all Anthropic was only cooking was Fable 5.1 and Opus 5? No way; they totally have unreleased models that aren't ready for the general public.

A lot of open-weight models (A lot, not ALL) have been distilled from current labs' frontiers. They aren't necessarily making breakthroughs in the same way frontier labs are. (I'm not saying they're making NO breakthroughs, but the biggest intelligence breakthroughs have come from the Frontier labs)

This post also hinges on the majority of Chinese/Open Weight labs staying open-weight. I'm not sure that in 10 years, these labs will continue giving away their models for practically free.

1

u/Fit-Gas-5760 8h ago

When did I said OpenAI or Anthropic will maie their models public? Wtf?

This post also hinges on the majority of Chinese/Open Weight labs staying open-weight. I'm not sure that in 10 years, these labs will continue giving away their models for practically free.

They def will, their business model is a whole different thing.

1

u/AkindaGood_programer 8h ago

Huh? What are you talking about? I never claimed that OpenAI or Anthropic would make their models public.

On your second comment: Their whole business model right now is based off open weight. What's gonna happen when we create so powerful models that the Chinese government doesn't allow them to give them away for free? Or maybe when... These Chinese companies actually want to start making money?

1

u/RepulsiveRaisin7 19h ago

People have been saying this about AWS for years but it never happened, not at scale. Don't get your hopes up.

0

u/UnexpectedFisting 12h ago

Lmfaoooo

Yeah I’m sure large enterprises will start migrating to local llm hosting anyday now. Wake me when you find a solution to hosting and running llm models the size of kimi k3 for a few thousand developers at nominal tps without spending roughly 100mil if you can even get the datacenter hardware through a supplier. And then you have the costs of hosting that hardware, insuring it, maintaining it, decommissioning over a time span, disaster recovery. I swear people here have the dumbest takes, hey why don’t we go back to local hosting compute infrastructure and company servers while we’re at it! I hear that’s cheaper too!

1

u/Fit-Gas-5760 10h ago

A thousand devs doing full agentic coding on top of an API will easily hit 100mil in less than a year lol do you actually know API prices and token usage in agentic coding? Also, a company with 1000 devs will pay way more than 100mil a year in salaries, its base math...