r/codex • u/OkMeat6773 • 14h ago
Humor Remember When GPT Would Just… Do the Task?
Remember the old days when you could ask GPT for a change, and it would just search your files and do it? No overengineering, no overthinking.
Now it thinks for 30 seconds → runs 10 commands → thinks another 30 seconds → creates 1–3 QA files and a bunch of SHAs → changes 3–4 lines of code → double-checks everything → oops, compacting → starts reasoning again with the latest request just to double-check it all over again.
11
u/EyesOfAzula 13h ago
I think this is happening because they are preparing for Astra.
It happened with 5.5 -> 5.6
The current version gets stupid then the new version comes out
10
u/Frosty-Ad1071 13h ago
I'm fairly sure it's a tactic to make the new model feel smarter and bigger upgrade than it actually is. I'm a conspiracy nut anyway so theres that
5
u/gachigachi_ 12h ago
More likely that it has to do with different kinds of compute being reallocated. Not all chips are equal.
-2
u/habeebiii 5h ago
“Not all chips are equal”? The fuck does that have to do with compute being reallocated??
2
u/Genneth_Kriffin 12h ago
It's not even a conspiracy, I remember when Claude 3.7 was the hottest shit on the market,
and I swear for medium level coding tasks nothing has really changed in any substantial way even though it is constantly hyped up.Sure, the result and accuracy might be a tad better, but I think that is mostly for upper end complex stuff or long running tasks where they are doing stuff without guidance. For more direct and user guided tasks, it's the same shit we had years ago but they burn 10x the tokens at inflated prices to confuse themselves with bloated context from their own thinking process.
Genuinely, I think the real issue is a clear lack of specific models.
We keep getting these upgraded flagships that keeps getting packed with books from 1925, sloppy AI projects from other users, coding manuals in morse code and ancient Chinese etc.We need models that are designed bottom up for specific uses, that keep getting refined for that.
That, or some break trough in a cheap , fast model that is just very, very good at taking custom datasets and instructions and following it.
I was hoping a bit for Gemini Flash/Flash Lite, because it do be fast af - but It can't stop itself from going absolutely of the rails now and then.
2
u/Innerdaze2600 8h ago
They’ve reallocated vram / premium compute to the new model ahead of release.
My production runs on an RTX 5060 Ti 16Gb comfortably.
My dev runs on twin RTX PRO 4000 Blackwells.
It is what it is.
Worker bees don’t get premium compute, creatives do.
1
u/Supermax64 9h ago
It's being independently benchmarked against day 1 benchmarks from earlier models so that would be somewhat pointless
6
u/Significant_Whole722 13h ago
Im not someone to complain but i feel 5.6 sol went full brain dead as of last night
2
u/shady101852 12h ago
Last week*
1
u/Significant_Whole722 12h ago
What if we are the ones actually just getting more dumb and lazy though
2
u/shady101852 9h ago edited 9h ago
Doubt.
I went from being able to tell codex what I want to having to tell it to chatgpt, and have it write a technical prompt for codex otherwise the dumb ass LLM will do some BS, waste time, make incorrect changes or just not even do what I asked in the first place. I did not have this issue a month ago. I only use sol max.
1
u/ninernetneepneep 7h ago
Sol Max may be your problem. Don't choose a higher reasoning level than the task requires.
1
1
u/wintermute023 6h ago
You know it was fine until this evening, then it just quiet quit on me. Lots of “step 1 I’ll extract the data and step 2 I’ll categorise ready for the next task, continuing with step 1 now” and then just stop.
Then I ask it to continue and it reviews the files again and does the same.
It refuses to actually do any work. Just going round in circles with it apologising for stopping over and over.
5
u/lolcatsayz 10h ago
The 300k context is becoming a joke compared to Claude's 1M, especially for complex coding tasks. The amount of reasoning these models do drowns out most of the token usage before compaction. In before "300k is plenty, the intelligence is already more than anyone needs for any task remaining in humanity, you need to make your code base more component based and your prompts more specific " bros.
2
u/TopSeaworthiness1679 13h ago
Sol thinks way too much and it is hard to change the course if it goes into wrong way. So it takes 2x more time and 2x more tokens to actually get things done than claude…. Claude at least is very fast at searching codebase. Codex sub agents take forever.
2
1
1
1
1
u/jonydevidson 9h ago
In the time it took you to cry here, you could've said this to Codex and ask it to create a /nobullshit skill that YOLOs the change without being thorough or verifying.
So whenever you want it to work that way, just invoke /nobullshit
1
u/Innerdaze2600 8h ago
Start processing business data bro, then you will really cry when you encounter the concept of ’providence’.
1
17
u/RealSlyck 14h ago
Wait, there was a time before it ever mentioned SHA? Thought I was hallucinating…now it’s just the models ☹️