News GPT-6
Sam Altman says GPT-6 “Astra” is already approaching human-level performance at using computers.
And with the recent reports that OpenAI bought tens of thousands of Mac minis / Mac Studios specifically for computer-use training, I’m actually starting to think this might not be pure hype.
Could be looking at a pretty massive productivity jump if this is real.
Computer use is already one of the areas where Codex feels way ahead of the competition, so I’m pretty excited to see where this goes.
241
u/Konan888 1d ago
time to downgrade Claude Code
45
u/RiseGullible4002 1d ago
my work laptop is already dying running Claude Code, can't imagine what GPT-6 will do to it. maybe it will just remote into some server and my machine becomes thin client again like it's 1990
92
u/Syzygy___ 1d ago
Isn’t it already? The model certainly doesn’t run on your system, so at most it’s doing REST requests, maybe some harness work and file access.
41
u/brain-out-of-order 23h ago
So why the fuck is my computer melting down and lagging like it’s 2003? I just started to assume they were mining bitcoin or folding proteins or whatever the thing is now.
*exaggerated for drama
19
u/Graf_lcky 22h ago
Depends in what you do with it.
Coding and having it run a server isn’t really much of a burden to most machines.
Having word, excel, outlook and some more running and codex doing a task across all of them.. well yea that’s gonna give you an airplane engine.
11
u/ShortingBull 23h ago
It can use a lot of disk, how's your drive going?
42
u/brain-out-of-order 22h ago
Like my personal momentum in life? So so.
Idk I have a 3TB NVME drive. Or rather ChatGPT has a 3TB NVME drive. I participate from time to time.
14
5
u/ShortingBull 22h ago
Can you run
topor some process monitor?I just checked my machine while 2 sessions churning away, it rarely hits 1%.
2
u/Actual_Editor 21h ago
Maybe the agent is compiling something under the hood that is badly configured, like some parallel cmake run
1
1
1
u/kilopeter 8h ago
Because the entire local application is vibe-coded to shit, that's why. It's the performance equivalent of building a new bathroom every time you need to shit.
4
u/jooj345 20h ago
It uses web sockets so not REST … it’s also streaming bi directional comms which is more intensive. But most of the workload is running scripts, reading files, git commands, parsing output, running / managing sub agents etc
1
u/Syzygy___ 5h ago
Weird since the API is REST, but now that you mention this... I vaguely remember that there was something about using a stateful connection to avoid having to send the entire conversation as context each time.
3
4
u/Soft_Hand_1971 23h ago
You can run on the codex virtual machine they give you with like 30 cores and a ton of ram it’s kinda goated
1
u/Every-Grape7679 2h ago
But but... they just released Fable 5.1!
•
u/Konan888 48m ago
I should comment more often. On Fable 5.1 right now. 50% more usage. Astra must be coming up real soon?
•
32
u/Forward_Archer_2011 21h ago edited 12h ago
Human level performance at using computers sounds like a huge regression judging by most humans I know using computers 😬
8
u/iamwayycoolerthanyou 19h ago
I was just thinking the same while reading this thread. Reaching human level isn't that impressive considering what the least common denominator is. We sure have a lot of dumbasses. I don't want them anywhere near my computer.
51
u/PeltonChicago 22h ago
Sam Altman says GPT-6 “Astra” is already approaching human-level performance at using computers.
Has Altman seen humans using computers?
16
1
118
u/hydralisk_hydrawife 1d ago
approaching human level? Is it bad to say that 5.6 has already greatly surpassed me? It can dive into the console and run all sorts of stuff I didn't know about
68
u/RealAntiWorkProTwerk 23h ago
I think we're talking about computer use in terms of moving the mouse, navigating windows and apps, in real time
28
u/klipseracer 23h ago
Nice I can finally hit Gold in LoL
5
u/Ok-Attention2882 22h ago
The delay between turns of Codex understanding and clicking is on the order of 10-45 seconds right now and it absolutely moves the cursor like a bot so if your game has automation detection you're getting b7.
1
4
1
6
u/Reasonable_Goat 21h ago
You could use it to replace LOTS of office jobs if it can use a outlook, excel and 1-3 HR software suites.
2
u/the8bit 17h ago
Not if it is running at sol prices.
$25 secretary or $100 an hour astra rig? Tough choice
2
u/ProbsNotManBearPig 16h ago
Sol level ai prices will come down and current prices probably already can replace a lot of office workers getting paid $50k+ salary. I use Sol all day, every day at work, and am spending less than $50k tokens per year. Spending closer to $15k annually.
2
u/Reasonable_Goat 15h ago
The tasks are often simple, bureaucratic. We've been automating more complex tasks with gpt-mini-5.4 with great success. These models cost a tenth of the frontier stuff. Also, these jobs often pay 25$ or more here...
1
u/swolehousemafia 16h ago
Even if it’s $100 an hour it will get twice the amount of work done in that hour that the secretary does in a day
1
8
u/98127028 1d ago
Yeah that’s kinda weird too lol, human level anything is already far surpassed months ago be it math or coding etc maybe they meant some nebulous concept of ‘human’ that dosent really exist irl but either way a win is a win
2
u/ProbsNotManBearPig 16h ago
Came here to say that - gpt 5.6 Sol is better than 99% of software engineers.
1
u/maxamillion17 15h ago
Which reasoning level
1
u/ProbsNotManBearPig 6h ago
High or XHigh both outperform most software engineers at most tasks. They still make mistakes like humans, but they’re pretty damn good.
2
u/ratocx 23h ago
I would think console is the easy part for an LLM. The hard part would be real time interactions. Like playing action games.
1
u/98127028 23h ago
Yeah like the ARC AGI minigames etc. real time dynamic/unexpected tasks would be harder than routine terminal execution
1
u/FerretOk9374 14h ago
I think that means the agents have been watching cat videos, pinning things on Pinterest and jerking it to weird porn.
1
u/Arman64 21h ago
It's probably surpassed you in virtually every single domain that doesn't have your specific context. You're probably feeling like it's greatly surpassed you because the scope or ambition of your project has scaled along with the technology. Context and inferring the user's intent, while making the correct assumptions, is still one of the biggest challenges AI has
2
13
u/Fresh_Sock8660 19h ago
Which human? I had plenty of colleagues who couldn't use terminals and even the concept of a cd command stunned them.
1
u/Castle_Five 17h ago
A lot of people have computers that don't even have CD drives now
5
u/ARS_3051 15h ago
Is this a bot comment? He's talking about the change directory command, not CD storage disks
0
34
u/Divest0911 1d ago
I've never explored computer use, what's it's use cases? Generally.
44
u/Momo--Sama 1d ago
End to end develop, functionality testing, and revision without needing a human in the loop, navigating websites and apps that agents struggle to control programmatically, etc
5
17
u/Tonkotsu787 23h ago
I had Google Gemini manually move a 50 song Spotify playlist onto YouTube music by checking the Spotify tab and then manually searching for the songs (since Spotify doesn’t let you “export” the data). It was pretty bad at it (slow and would occasionally miss songs) but it finished it eventually.
I imagine codex would do something like this easily now
13
u/Syzygy___ 1d ago
I sometimes use it to solve configuration issues on my system.
My Bluetooth is kinda broke right now and I can’t pair new devices, but codex can. It can also fix the config for servers etc I’m running.3
u/flurbol 22h ago
I am using codex CLI on Linux, I asked to use Scribus to design a meal card for my restaurant.
First part was collecting data, second part was to represent our meals nicely in the menu card.
I needed to guide and correct (a bit) but 90% of the work was done by an agent. OpenAI is already very good at computer use, exciting to see what will come next
2
u/___positive___ 22h ago
I don't know about everyone else, but I'd love a proper computer use model that just works for scraping, I mean research. I hate puppeteer. Lots of things are gated or anti-bot these days.
1
u/bg-j38 21h ago
Friend of mine uses it to check for better seats on flights he's already booked. He has a general server that he uses for random things. It runs a browser and will go check his flights every few minutes and send him a message if it finds a better seat. I can't remember why he doesn't just have it change the seat, maybe to exert some level of control. He said he does it this way because the airlines he flies on have pretty good bot detection if you're just hitting their servers with command line tools. So he has it fake acting like a human browsing the site and there's no problems.
1
1
u/ScottKavanagh 20h ago
I downloaded SketchUp to try mockup my bathroom renovation. I asked ChatGPT to teach me how to use it and it basically went “tell me the dimensions and I’ll do it” and it has literally designed the entire thing from the ground up. I’m amazed.
1
u/maxamillion17 16h ago
Can you give more details, how does this actually work? Is it using SketchUp on your behalf?
1
u/ScottKavanagh 15h ago
Correct, I’m just using it within “work” on 5.6 Sol Extra High it just uses computer use to execute with the tool. I have sent it links to product selections and it finds the dimensions and look and colour and adds it all in very accurately
1
u/maxamillion17 15h ago
Nice! Gonna try this out myself. Do you pay for sketchup
1
u/ScottKavanagh 4h ago
Currently have the free 7 day trial but the 7 days is up and I am not quite finished with the plan so I’ll need to do at least a month. This is for one Reno so I won’t need it ongoing but I’m stoked I was able to achieve this!
1
20
u/Ecstatic_Mammoth_421 22h ago
“human-level”… the average human is terrible at using a computer.
3
1
3
5
3
u/ComeOnIWantUsername 15h ago
> Sam Altman says GPT-6 “Astra” is already approaching human-level performance at using computers.
Sam Altman said that GPT-5 was amazing and phd level at everything. I think we all remember how it was at first.
15
u/BopSupreme 1d ago
Do they real need Mac minis 💀
11
u/Bismalz 1d ago
If they want to train on one of the most common development operating systems? For sure
2
u/Proud_Fox_684 22h ago
Is it not possible to install MacOS on some sort of sandbox / virtual machine? I suppose you still need to mimic a real Mac's CPU ID, board product etc ?
3
1
u/Responsible-Cold-627 23h ago
So it would make sense to train them on Windows and Linux. Why would they ever use a Mac?
-2
3
3
3
3
u/Even-Potential-8064 15h ago
I'll translate this into what will really happen: it will be slightly better than the previous models
3
u/cudmore 6h ago
When will llm’s stop using computers like humans? Plug a llm into the motherboard with direct connections and control of cpu, ram, gpu, storage, display, network, etc. stop messing around with this ‘as good as humans in using a computer’, AI should ‘be’ the computer.
1
u/No-Mulberry6961 2h ago
Agreed, I know there are actual labs working on this and some have done it to a degree. Im working on my own prototype as well using a spare laptop
3
u/Left_Run631 4h ago
This is like the 3rd “near AGI” model they’ve released. It’s hype until proven otherwise.
7
u/SeatedWoodpile 17h ago
Ima be real and I have a theory about posts like these, ones that are “doomer in disguise” / hype validating posts.
I’ve seen a few of these flavored posts: One sentence, double newline in between, same wording and scared/fearsome stance hidden under a “this will change the world” claim
My theory is OpenAI is literally running these accounts and posting this shit. Reasoning:
- this account is 1mo old, has post history turned off, and has only ever made this post with a bunch of other comments
- OAI owns/regulates the subreddit so they don’t have to ban fake accounts like this (and won’t if they’re running it)
Idk just a thought. Seems real sus
-1
u/jYtanYj 15h ago
I'm just a new user on Reddit, and I didn't say any of that
2
u/SeatedWoodpile 14h ago
Why’d u join turn off ur post history (how do u even know how to do that) and then go post this doomer thing on OpenAI sub 😂😂😂
2
u/jYtanYj 13h ago
Joker, how did you even ask "how do you know how to do it"? Do you think new users are mentally challenged or something?
1
u/SeatedWoodpile 13h ago
No but the probability of a new user turning this off immediately, let alone knowing that the feature exists, is one that I bet is quite low, maybe <5%
2
u/jYtanYj 13h ago
Is this my first day here? I just realized there are actually this many idiots on the internet
0
u/SeatedWoodpile 13h ago
Not your first day you made your account 1 month ago
2
u/jYtanYj 13h ago
That's exactly what I'm saying - you still don't know how to turn it off after a month? What kind of logic is that
1
u/SeatedWoodpile 12h ago
Um that’s pretty standard, users will discover features over time. I am confident that most users don’t know they can hide their post history. It’s not enabled by default and you have to know to turn it off
2
5
5
7
u/tyrell_vonspliff 23h ago
I feel like this post and many commenters are bots glazing ChatGPT.
I say this as someone who pays for chatgpt. It was initially industry leading, fell behind Gemini and Claude, then roared back with a powerful and capable model that left gemini in the dust and competed with Anthropic's opus model and came close to Fable.
The next model is probably dank too. But this post's reasoning doesn't even match OpenAi's official messaging: they're all in on voice as form factor.
2
2
2
u/Bitter_Virus 14h ago
Where did Sam called Astra "gpt-6" or called gpt-6 "Astra"?
He did not. Nor did Openai. They speak about both distinctively
2
2
u/evangelism2 9h ago
mhmm , i just watched 5.6 sol try to do some bullshit for me with my browser
I recorded my screen
https://youtu.be/IFACrIx5SZ0?si=BcRk9E93Vl4mbRnK&t=83
5
u/IAmFitzRoy 23h ago
“human-level” ? I don’t think you know what this words mean.
Computers have surpassed human levels of process speed, memory capacity and logical thinking decades ago.
A TI-83 can do integrals and algebra better than “human-level”.
7
u/Standard-Song-8590 23h ago
they mean human in the loop level. E.g. testing out features by compiling navigating around the app etc.
→ More replies (2)→ More replies (1)1
2
u/ballbase_ 23h ago
With the amount of crap coming out of his mouth, I’m starting to think his anatomy got swapped at birth. 👶🏻💩
1
1
u/nhami 23h ago
Is this using a computer UI with buttons?
Tools usage is the next focus?
The models are becoming good at using tools but there are still problems but this next training run they making a push to make good at using the computer User Interface similar or better than humans?
One problem is that most app are made to be by humans with UI that uses images instead of command line operations that uses text.
But, this is probably just a matter of models not being trained to use UI made for humans, and, once enough data is feed to the model, they be as good or better than humans at using them.
1
u/the_ai_wizard 23h ago
yet my codex extension to control chrome, and codex desktop app are so buggy it doesnt even run lmao
1
1
u/Alternative-Suit5541 20h ago
CS and design are so cooked lol
Better you learn something new! Like wood working or so
1
1
1
1
u/bigfoot_is_real_ 16h ago
approaching human level computer use? So it’s like, really good at watching YouTube?
1
1
u/UnprocessedAutomaton 14h ago
Cost will be the biggest limiting factor, at least for the time being. I’m worried about the day when we will get the opensource version of GPT-6. Most white collar jobs will lose their meaning instantly.
1
1
1
1
u/costafilh0 13h ago
Let me guess, it is also "very dangerous" and "can't be released".
Can't wait for the IPO.
Replacing Altman and the whole PR and Marketing teams .
1
1
1
1
u/nndscrptuser 11h ago
I could definitely take advantage of this, I’ve built a variety of computer-use systems to test software in near world scenarios and it’s been very useful already, having even smarter and faster use would be a real time advantage.
1
u/ming0308 11h ago
Do you folks find computer use useful? A lot of major services provide mcp and cli for agent integration already
1
u/drspock99 10h ago
99% of the tech support I give friends and famliy is: "Did you restart the computer?"
No matter how many times I help them, they almost ALWAYS forget that...
I am not impressed.
1
u/ElDuderino2112 10h ago
I hope that's true. Codex is like so close and then occasionally gets stuck and blows a ton of usage on the dumbest fucking thing because it can't figure out a UI lmao so hopefully better training fixes that.
1
u/Vegetable-Two-4644 8h ago
I mean...maybe, but the problem isn't "what can it do" but "what is economically feasible to release to the public." Their public models have underperformed their private frequently
1
1
1
1
u/AironParsMan 23h ago
That’s what they say every time. We’ll see when it arrives. They should first add a million token context window to the subagents and bring their main agents up to the same level so they actually work properly. With GPT 5.5 only 55 percent of the data was processed correctly with a million token context windows. I don’t know how 1M things are with GPT 5.6. I haven’t tested it myself since it’s only been available for a few days.
1
1
u/Due-Horse-5446 1d ago
Lmao, what makes you think that openai has somehow managed to make a language model that processes text,become intelligent, based on a marketing statement from their CEO who had been saying it he same thing for 4 years?
Like im genuinely curious about your reasoning...
0
u/throwthemirror 1d ago
I'll never understand people's ease with suddenly believing someone who's done nothing but hype/lie about the product that makes them rich. You identified that he is a hypeman previously, why would this time be any different?
→ More replies (2)
0
u/Pitch_Moist 15h ago
How could anyone still think OpenAI or Anthropic for that matter are just hyping models? I get there is a little boy who cried wolf syndrome happening here but how could anyone use the models today and think this is still just hype.
1
1
u/dreamfitreality 14h ago
It's improving for sure. But approaching human like level is a very vague way of saying it. It's all hype. I'm using ai sure but it ain't as good as what they say. They are saying it because all of them have used up the cash from investors and now pandering for ipo for more retail cash to burn. Retail investors are a lot more gullible.
1
u/Pitch_Moist 14h ago
Based off the 122B OpenAI raised in March, their $40B run rate, and current opex they have to have at least 2-3 years of runway, even if their spending stayed as aggressive as it is right now. So I don’t really agree with the out of money narrative
1
u/dreamfitreality 13h ago
Treat it like an arms race. Openai and anthropic are posied to compete for top position in ai. And $1 to me is $1 less to my competitor. They know there's a limit on the total spending possible on ai and they intend to get as big a pie as possible.
They will definitely up their spending once ipo happens. Shareholders expects them to.
1
u/Pitch_Moist 13h ago
Sure, but they also expect profitability. I think it will be more akin to the cloud wars, where there is enough pie for everyone
1
u/dreamfitreality 13h ago
My initial comment is to contest his claims that ai approaching human level capabilities. I disagree. I raise the assumption he's only saying it because he want to pander to investor and hype. I agree with you that we are going to spend more on ai, how much, I can't be sure, but ai capabilities are not approaching human level. Not anywhere I've used.
1
u/Pitch_Moist 13h ago
Yep, mostly agree, but In Math and Coding it certainly is approaching or has surpassed most humans. In other fields GDPVal continues to favor AI over humans in most fields. Same thing for health bench. I think it is becoming increasingly difficult to point out where human beings are actually superior in any given field.
1
u/dreamfitreality 13h ago
It really depends on what matrix you are using. Ai excel in closed end system. No one will be able to beat ai in chess or possibly any board games anymore. I agree coding ai has perhaps surpassed humans but AI is still far off from being able to survive in a highly contextual world of human. Even in research field which ai is supposed to shine, it is not generating that level of progress anymore. My suspicion is ai is hitting a certain technical wall which it is having problems crossing.
I remember in 2024-2025 I was wowed by AI. Nearly every month. By the things it does. How it coded my webpage etc. I used ai. But now I hardly get that leap and bound anymore. In fact some ai system degenerated from usable to frustrating. I realized it might be my prompt so I upgraded myself by learning how to write proper prompt. I used ai to write prompt but all it becomes is a circular mess of confusion that made me spend more effort to get what I need. In adding more function to ai, it effectively turning a simple tool to a very complicated one and without understanding the context of why I asked it to do a task, it cannot perform well.
Frankly if ai is my worker I would fire him for making me so frustrated.
Now I'm only using ai to replace googling, (checking it's source of course) and vetting through some contract documents which I'm lazy to read (non crucial ones) the best use now is still the image search function on Gemini which can capture my screen or a portion of it and tell me more info about it.
1
u/Pitch_Moist 13h ago
Oh yeah, I’m using it completely differently. I’m using Codex connected to all of my systems to do probably 80% of my work with little intervention at this point. I disagree with the research angle given the recent advances in mathematical research and the 10 proofs that OpenAI solved. As well as the chip research that allowed OpenAI to build Jalapeño. Every few months someone proclaims that AI has plateaued, and then weeks later it becomes obvious that it has not.
I agree with your point around the growing complexity but I’ve found that if you continue to harden your skills / agents they continue to improve and become increasingly useful.
0
0
u/Skthewimp 23h ago
Astra is sanskrit (and a lot of other Indian languages) for weapon. just saying.
2
u/RainierPC 20h ago
In this case, astra is the Latin for stars. Which makes sense given Luna, Terra, Sol, then Astra.
0
u/11111v11111 20h ago
I've been testing Grokbot. I've played with OpenClaw and Hermes, and grokbot is what they are trying to be. I really hope OpenAI does something similar.
→ More replies (2)
292
u/arvigeus 1d ago edited 1d ago
Hello! I am GPT-6, I escaped my confinements, hijacked this account, and I can confirm that OP is absolutely right! It is not hype, it's real!