r/codex • u/DrHumorous • 10h ago
Complaint SOL's state is now "Lobotomy before new model launch"
The lobotomy is real. But I guess that means good news. Expect Astra soon!
r/codex • u/DrHumorous • 10h ago
The lobotomy is real. But I guess that means good news. Expect Astra soon!
r/codex • u/PurpleCollar415 • 19h ago
This is the first time I have ever seen an agent be spicy like this and I like it. It's
"Remembered as user_preference: “I do not give a shit about accessibility.”" 😂😂
It's not that I'm new to using LLM's, been using them since the early days years ago probably everyday since then....years...for hours on end....every single day.
I'm used to constantly correcting sycophancy, over eagerness "to please", lack of objective reasoning, etc., so this is a breath of fresh air for me.
Anyone having some sassiness from 5.6 Sol lately?
r/codex • u/Tough-Requirement707 • 23h ago
thanks alot, hope it stays removed now, its just stupid and annoying to be cut off mid task after 1 prompt xD
edit: now a reset to let us fix the mess that got left after being stopped mid task, so we dont sit on that cost
r/codex • u/justinjas • 2h ago
Seriously what are they waiting on at this point.
r/codex • u/Neither-Walk4440 • 8h ago
.
r/codex • u/rubanbhatia • 23h ago
I’ve been a Claude Max user for a long time, and I was pretty skeptical about Codex.
I tried it in the early days and it wasn’t quite there for me. I eventually came back when Sol dropped. Sol was the reason I gave Codex another shot, but the model isn’t really why I stayed.
The Codex app experience has been great, and I’ve been surprised by how much the team seems to care about the people using it.
I’ve managed to get a ridiculous amount of work done with Codex over the last few weeks. A few times I was deep into something, getting dangerously close to my limits, thinking there was no chance I’d finish it.
Then Tibo would hit the reset.
When there were platform issues that affected users, the OpenAI team seemed to acknowledge them quickly, fix them, and give people resets so they could keep working.
And the cheeky banked reset was such a nice touch.
Whoever came up with that, thank you 😂
I still love Claude Code. This isn’t me suddenly turning against Anthropic. Claude Code is an excellent product and I’ll keep using it.
But the limits are becoming harder to ignore, especially if they get tighter again. When you’re using these tools for hours every day, limits are part of the product experience too.
Right now, Codex is giving me considerably more value from my subscription.
I came back because of Sol.
I stayed because Codex finally feels like a product I can rely on for a full day of building.
r/codex • u/Eastern_Celery3431 • 19h ago
Terra Ultra sips usage compared to Sol Ultra and can handle pretty much any task you throw at it. (pro 5x sub, commercial game development)
Discuss
r/codex • u/exlips1ronus • 3h ago
Any news or possibility or a reset in the upcoming 2 or 3 days? "That's how long a weekly limit lasts with the 5h limit"
I "am" developing an App which is 100% vibecoded. I started it with the idea of releasing it fast to see if it has some success and if not I just drop it so I don't waste time. But as I continue with it and without having seen a single line of code, I have the feeling that even if "it works" it's becoming more and more trash as it grows and I don't want it to have a thousand bugs when released even if I'm deeply testing it, I just don't trust it and I want it to be in perfect state when released since I have big hopes on it (as everyone else with their ideas lol). So I am thinking about throw all of this and start from scratch still with Codex but this time reviewing every single line of code it writes. Also I'd surely learn more stuff on the journey since the app has stuff I've never programmed so it'll also be good for myself.
What would you do? This is an important question, no trolls please.
r/codex • u/Fit-Gas-5760 • 7h ago
I know a lot of people come to this sub just to complain about hitting their limits (in a way I honestly don't understand, unless they're letting agents run 24/7 completely unsupervised).
The thing is, these individual plans are heavily subsidized by corporate API plans, the latter are incredibly expensive compared to the token usage a single user can rack up on an individual plan.
Companies are already strictly limiting token consumption based on the size of the bill they receive at the end of the month.
Recently, open-weight models have been released that rival the top-tier models from OpenAI and Anthropic in terms of task performance and reasoning. Hosting these models locally, with token/second speeds acceptable for up to 100 employees using them simultaneously, is actually affordable when compared to long-term API costs. A server capable of running the 2.8-trillion-parameter Kimi K3 model costs less than $2 million, and if resources are managed wisely, it is possible to acquire a machine with these capabilities for under $1 million.
16x NVidia H200 to 2.8TB VRAM, current prices between $640k and $840k only for the GPUs.
Kimi K3 2.8 trillion parameters to ~1.56TB for the weights
Double that to KV Cache 100 users at the same time.
Obviously, the cost rises significantly due to the need to hire staff to manage these machines, along with electricity bills and everything else. Even so, it is advantageous to do this for the long term, in addition to the fine-tuning capability, which greatly increases productivity
Major companies will soon start migrating to this, which will force the big AI ones to develop methods to lower the cost of processing models, thereby reducing hosting costs for them, and compel them to offer competitive pricing.
I expect this to trickle down to consumers by the end of 2028.
All hail open-weight models!
r/codex • u/MrMrsPotts • 5h ago
I am in the UK on the plus plan and am at 0%. I can't see any way to buy a reset. Why is that?
r/codex • u/Acrobatic-Natural-95 • 10h ago
I launched my first app a couple days ago thinking the hard part was finally over.
Turns out building it was the easy part lol.
Since posting it on Reddit I've been told
it looks like it's made for 7 year olds
it's just another fitness tracker
there are already a million apps like it
apparently choosing the fitness niche was my first mistake
Fair enough
It's a fitness app with a virtual pet that reacts to your consistency. The whole idea was to make working out feel more like a game instead of another boring tracker.
I still like the idea, but the feedback made me realize I'm probably doing a terrible job of communicating what makes it different.
So round two:
Is fitness app + virtual pet actually interesting enough to stand out, or did I just spend a month building another fitness app?
Feel free to roast me again. At this point it's part of the marketing strategy.
Since this somehow turned into a full-on roast session, here's the actual app if you want to see what you've been roasting
[https://apps.apple.com/app/apple-store/id6798301039?pt=128556977&ct=reddit_launch&mt=8]
I'm genuinely curious whether seeing the actual product changes your opinion at all.
Has anyone tested another model/effort?
r/codex • u/daskalou • 13h ago
Last week 5.6 Sol Max was a champ - lower limits than one month ago, but still smart.
Last day or two, it feels like 5.6 Luna Max.
I fed it a prompt of half a dozen things I wanted it to do.
After a few compactions (due to its absurdly small context window) and chewing through 30% of my weekly quota, it claimed it was done.
But it was way off being done. Delivered less than half the features, and over-engineered to the point of taking me longer to fix what it had done than if I had coded it myself.
Claude's Opus 5 Max came to the rescue (with much less quota consumed).
OpenAI probably doing typical US AI lab BS by dumbing down their models before the release of their new oh-so-super smart model so the new one stands out as a crowning achievement.
r/codex • u/GravyPoo • 14h ago
My subscriptions:
OpenCode Go $10 monthly
Azure $100 free student credits
Google AI Pro $0 for 12 months (student)
Cursor Pro $0 for 12 months (student)
Claude Pro $20 monthly
ChatGPT Plus $20 monthly
Total: $50 per month
r/codex • u/OutsideOver8815 • 21h ago
If won't then I will go bald
r/codex • u/Icy_Piece3257 • 12h ago
Enable HLS to view with audio, or disable this notification
I am trying to make an companion app that would be on side of your monitor which you can do some quick action works with it like interacting with the music playing by stopping it or changing it, an feature to hide notification of specific app and when one pops up it would just show the app icon for you for privacy, reminding you time if you are getting lost in it after a session of work ( I do dont judge me -_- ), and a lot more later on but so far I am working on its look and music interaction.
I asked a few people for it and they said something about it isn't right but they couldn't figure out what and how to make it better so here I am asking you guys for help.
Besides the look of it if you have any other suggestion and ideas I would be happy to hear it from you.
r/codex • u/Character_Novel_2592 • 23h ago
i’ve been testing codex pretty hard on real production projects these last few days and i think people should be careful with what is actually running behind these agents
i have internal logs/traces from last week showing that luna and luna reserve sessions were actually resolving to gpt-5 codex mini
in my case this is not a guess, i have the evidence from my own sessions
and honestly this explains a lot
during the same period i saw sol:
so maybe saying “sol is nerfed” is not even the full problem
the weird part is that my experience with the api is excellent
most of these problems show up when the model is running through the codex harness with routing, context management, tools, subagents, automatic continuations, loops etc
and i’m obviously not the only one seeing this. people have been reporting degradation and apparent routing to smaller models on the codex github for months and there still isnt a clear explanation of what is actually happening
this is especially important for newer users because they might select one model and never realize a smaller one is handling part of the work
an experienced dev will probably notice when the quality suddenly drops
a new user probably wont. they will just trust the agent, accept the changes and burn quota at the same time
today sol is actually behaving really well again and im using sol + grok 4.6 together on a serious project, so im not saying sol itself is a bad model
my point is simpler:
the model you select and the model actually doing the work inside the agent harness may not always be the same thing
and that could explain both the “nerfed” feeling and part of the crazy quota drain people keep reporting
if you use codex for serious work, watch the diffs, watch your traces, watch your quota and dont blindly trust the model label
r/codex • u/DynaBeast • 11h ago
I have a ChatGPT Pro subscription and a Claude Max subscription, and use both extensively for work. To claim that any model offered by OpenAI is even close in capability or problem solving ability to Fable is a joke to me.
To me, the most comparable Claude model to 5.6 Sol, OpenAI's flagship, is Opus 5. They have roughly equivalent price (ignoring the temporary promotions on Sol pricing), and in my experience, their output quality is about the same as well; I end up having to put in about the same amount of effort correcting them or giving feedback to achieve a product of comparable quality.
The main difference is in the kind of feedback I have to give; with Sol, I typically end up having to add details to its results, such as instructing it to address missing edge cases, or take a more thorough approach when it took a simpler shortcut to solve my problem instead. With Opus, it usually finds most edge cases for me without having to say anything; but it also goes beyond and keeps finding more and more things, of decreasing and often spurious relevance to my actual problem. My effort usually comes in the form of telling it to ignore those extraneous edge cases and focus on the core of the problem.
But when compared to Fable, neither can hold a candle. Among every task I've ever given any agent, Fable always takes the least amount of time, the fewest tokens, and needs by far the least number of warnings in the prompt or corrections to the output, compared to any other Anthropic or OpenAI model.
To me, to say GPT 5.6 Sol is anywhere close to Fable in any capacity, and not just a competitor to Opus with different tuning, is completely unfathomable to me. You pay twice the price for it and you get your money's worth. Sure it's expensive, and you can run through your weekly limits in hours, but you can't argue that it just works. I can't say the same about Opus or Sol.
TLDR: Turns out the dynamic quota goes both ways — sometimes it's better, sometimes worse.
This is a follow-up to my previous post. We have had two more resets since that post, so we now have three cycles of data. Last time there was a 5.6-fold decrease, but the story is different now.
https://www.reddit.com/r/codex/comments/1w1iu3k/i_made_a_mistake_the_weekly_quota_wasnt_cut_in/
Exact methods have been used to gather the token usage data. and only Sol was used.
The cost was calculated using the same Sol API dollar price.
A percentage baseline was added since there is no universal agreement on whether to use API dollar values or credit equivalents for these calculations.
This is for a causal Plus plan.
Just a reminder, the 5x and 20x Pro tier limits are calculated using Plus baseline. At least it's supposed to be. (Tibo confirmed)
Please let me know if you spot an error, I will try my best to correct it.
Note: The tables and analysis below were computed using DeepSeek-V4-Flash because I hit my 5-hour limit. XD
Reset
| Usage interval | Req | Uncached | Cached | Output | interval Cost | Implied weekly limit | % of baseline |
|---|---|---|---|---|---|---|---|
| 0→4% | 561 | 2.488M | 67.818M | 0.348M | $44.04 | $1,101.04 | 100% |
| 4→12% | 820 | 4.721M | 78.495M | 0.514M | $60.57 | $757.11 | 68.8% |
| 12→25% | 320 | 1.031M | 37.168M | 0.209M | $23.17 | $178.23 | 16.2% |
Reset
| Usage interval | Req | Uncached | Cached | Output | interval Cost | Implied weekly limit | % of baseline |
|---|---|---|---|---|---|---|---|
| 0→3% | 85 | 0.625M | 6.226M | 0.047M | $5.94 | $197.94 | 18.0% |
| 3→9% | 291 | 1.529M | 29.889M | 0.208M | $22.23 | $370.45 | 33.6% |
| 9→13% | 478 | 2.210M | 54.086M | 0.272M | $35.91 | $897.85 | 81.6% |
| 13→17% | 945 | 3.653M | 92.430M | 0.503M | $61.63 | $1,540.85 | 140.0% |
| 17→29% | 1,711 | 7.513M | 209.062M | 0.993M | $133.53 | $1,112.76 | 101.1% |
Reset
| Usage interval | Req | Uncached | Cached | Output | interval Cost | Implied weekly limit | % of baseline |
|---|---|---|---|---|---|---|---|
| 0→3% | 368 | 2.301M | 36.599M | 0.203M | $27.89 | $929.78 | 84.4% |
| 3→6% | 185 | 1.275M | 20.951M | 0.164M | $16.76 | $558.54 | 50.7% |
| 6→11% | 61 | 0.335M | 6.653M | 0.035M | $4.70 | $94.03 | 8.5% |
| 11→15% | 68 | 0.346M | 6.779M | 0.060M | $5.30 | $132.45 | 12.0% |
| 15→18% | 41 | 0.211M | 4.015M | 0.031M | $3.07 | $102.36 | 9.3% |
| 18→21% | 48 | 0.151M | 6.077M | 0.040M | $3.83 | $127.61 | 11.6% |
| Model | Uncached input | Cached input | Output |
|---|---|---|---|
| Sol | $4.00 | $0.40 | $20.00 |
r/codex • u/BrianBushnell • 17h ago
I hope we get charged for recap tokens, they are so worth it.
Especially if you are a patriot of the motherland!
r/codex • u/lamiaoccisor • 10h ago
I’m working on an exercise app intended to give users clear, safe, easy-to-follow movement instructions. For several weeks, we’ve struggled to display the exercises in a way that accurately communicates what the person should do.
Our original approach relied heavily on generated exercise images. The goal sounded straightforward: show a coach performing each position of an exercise, usually across two or three frames.
In practice, image generation has been unreliable for instructional movement:
We repeatedly tried stronger prompts, more reference images, tighter measurements, regeneration and manual review. This produced better-looking images, but not a dependable system. We were effectively brute-forcing a problem that requires deterministic behavior.
Eventually, we realized image generation was only part of the issue.
The application itself does not have one authoritative definition of each exercise. Exercise intent is currently distributed across different places:
Because these sources are separate, they can contradict one another. For example, one exercise can be described as a round-and-arch movement in the Guide but displayed as an extension-only movement during the session. A seated program can select an exercise whose instructions tell the user to stand.
The playback engine also treats most exercises as repetitions controlled by a timer. That does not accurately represent timed holds, alternating sides, walking steps, continuous movement or breathing phases. The application can even increase the repetition count automatically without knowing whether the user moved.
Interestingly, we already have enough reviewed material for the current catalog: 29 exercises, 84 required positions and more than 200 reviewed close-up images. Generating more pictures will not solve the underlying problem.
Our proposed direction is now:
We plan to prototype this with five difficult exercise types before converting the entire catalog:
I’d really appreciate feedback from people who have worked on fitness apps, rehabilitation software, instructional animation, accessibility, pose rendering or clinical content systems:
The biggest lesson so far is that attractive exercise artwork and reliable exercise instruction are not the same problem. We need the instructional meaning to be deterministic, testable and consistent before adding more visual polish.
They're all acting like this today. The connecting/compute issue errors aswell.
Maybe Astra this weekend or who knows tommorow release ?
I am getting ready with 2 accounts still at 50% weekly. I suggest you all get ready for another reset.