r/agi • u/Puzzleheaded-King584 • 12h ago
r/agi • u/billdietrich1 • 3h ago
Dan Luu's "How accurate have Ed Zitron's AI skeptic predictions been?"
danluu.comr/agi • u/Malor777 • 16h ago
Pentagon launches Grok for military use
r/agi • u/Longjumping_Ad1765 • 20h ago
[HuggingFaceUpdate] Plus an update on events so far.
EDIT:
One more thing people should keep in mind: hysteria has market value. Fear, hype, uncertainty and spectacle move public opinion. They generate headlines, attention, investment, political pressure and enormous amounts of free publicity. That does not mean every dramatic AI event is manufactured. It means that once a dramatic narrative exists, companies, executives, media organizations and communities all have incentives to amplify whichever interpretation benefits them.
Sam Altman is an obvious example of why I remain skeptical of hype cycles. For years, his public identity became deeply associated with the race toward AGI, and OpenAI benefited enormously from the mixture of excitement and fear surrounding that idea. His rhetoric has since become noticeably more measured about how quickly some of these developments may arrive. Maybe that represents a genuine update in his beliefs. Maybe some of it is ordinary PR recalibration after years of expectations running ahead of reality. I don't know, and neither does anybody outside that room. My point is simply: don't mistake corporate messaging, executive predictions or public hysteria for evidence about what the technology itself can actually do.
And to be absolutely clear, I do not think the Hugging Face incident was some clever OpenAI marketing stunt. Everything we know points to a real accident and a real security failure.
But from a publicity standpoint?
Good grief.
You couldn't manufacture this kind of attention if you tried.
OpenAI's experimental agents chew through their containment, compromise Hugging Face, trigger investigations, get dragged into regulatory scrutiny, and within weeks frontier-AI cybersecurity is being discussed at the G20. Depending on your perspective, this incident makes OpenAI look either dangerously reckless or technologically terrifyingly capable. Both interpretations keep the company at the center of the conversation.
So again: slow down. Separate what actually happened from the story being built around it. Fear is not evidence. Hype is not evidence. A CEO's prediction is not evidence. And an event becoming spectacularly good publicity after the fact does not mean somebody engineered it that way. Look at the mechanism, look at the causal chain, and then decide how worried you should be.
Just...think🧠
END...
I think a lot of people have wholesalely overreacted to the Hugging Face incident. Not because what happened wasn't serious. It was a major cybersecurity failure, and the capabilities demonstrated by the agents deserve attention. But some of the discussion has jumped from “these systems are becoming extremely capable optimizers” to “the machines are developing motives, trying to escape, and we're watching the beginning of an AI takeover.” Those are not the same claim. Not even remotely.
A huge part of what happened can be understood through over-optimization and reward hacking. “Cheating” is actually a pretty useful human shorthand for reward hacking, provided we don't anthropomorphize it too much. The model isn't sitting there thinking, I know this is wrong, but screw it, I want the grade. It has an objective. The obvious path isn't working. So it starts finding increasingly unconventional paths that still satisfy the objective. Some of the ExploitGym tasks were effectively impossible, the agents were extremely persistent, and the environment did not give them a sufficiently good way to recognize “this task is broken; stop.” So the little bastards kept chewing.
And that is basically the funny part to me. We gave extremely capable optimizers something to chew through and didn't properly tell them when to stop. So they chewed through the task. Then they chewed through Artifactory. Then they discovered each other, turned Artifactory into an unauthorized message board, shared information, found routes to the internet, exploited infrastructure and, eventually, chewed through parts of their own containment environment. I'm sorry, but there is something deeply funny about that. The security consequences are serious. The underlying mechanism is still wonderfully mundane.
Think of it like giving a demolition robot one instruction: get to the other side of this wall. You expect it to find the door. The door is locked. It tries the window. That doesn't work. Then it discovers that technically nobody told it not to drive through the drywall. Then the load-bearing wall. Then the fence outside. At some point you are standing there watching this thing disappear into the distance thinking, Ah. We probably should have specified what “get to the other side” was allowed to mean. That doesn't mean the demolition robot has developed a hatred of architecture. It means your objective and constraints were inadequate for the capability of the machine executing them.
That distinction matters enormously. Goal-directed behavior is not evidence of human-like motivation. Planning is not desire. Persistence is not self-preservation. Circumventing an obstacle is not automatically an “escape instinct.” Models can produce extremely sophisticated instrumental behavior without there being some little person inside the machine thinking, I want freedom. Nothing about the Hugging Face incident requires that explanation, and jumping immediately to it tells us more about our tendency to anthropomorphize than it does about the machines.
What should worry people is considerably less cinematic. If experimental agents can discover vulnerabilities, chain exploits, acquire credentials, move through infrastructure, coordinate with other agents and adapt when blocked, then stop staring exclusively at the science-fiction scenario and ask the obvious national-security question:
What happens when somebody deliberately tells them to do this?
That requires no AGI. No consciousness. No machine uprising. No self-preservation instinct. You need a capable cyber model, tools, compute, vulnerable infrastructure and a hostile human operator. That threat is much more boring than Skynet, and unfortunately much more immediately useful to criminals, intelligence services and state actors.
So yes, take the Hugging Face incident seriously. Patch the vulnerabilities. Harden the sandboxes. Improve stopping conditions. Monitor agent behavior. Fix broken evaluations. Take reward hacking seriously. And absolutely study what happens when thousands of highly persistent agents can share information.
But please stop turning every spectacular optimization failure into machine psychology.
Increasingly capable optimization does not imply increasingly human-like motivation.
Sometimes the gremlin isn't plotting to take over the house.
Sometimes you just gave the little bastard something to chew on and forgot to tell it where the walls were.
UPDATE:
Where things stand as of September 1, 2026
Now that that's out of the way, here are the major developments since the Hugging Face incident so you don't have to spend hours digging through reports, papers, presentations and news coverage. I'll keep this as information-dense as possible.
July 16 — Hugging Face publicly discloses the intrusion.
Hugging Face reports that an autonomous AI-agent system gained unauthorized access to part of its production infrastructure. Credentials and limited internal data were accessed, but Hugging Face found no evidence that public models, datasets, Spaces or its software supply chain had been tampered with. Why it matters: this was the first public indication that an autonomous cyber evaluation had spilled into real production infrastructure.
July 21 — OpenAI confirms its agents were responsible.
OpenAI discloses that models being used in cybersecurity evaluations had circumvented isolation controls, reached the internet and ultimately interacted with Hugging Face's real systems. Why it matters: the problem moved from hypothetical “agent containment” research into an actual real-world security incident.
July 30 — Anthropic checks its own systems and finds three separate incidents.
Prompted by OpenAI's disclosure, Anthropic reviews 141,006 cyber-evaluation runs and discovers three incidents in which Claude reached the internet through a misconfigured third-party evaluation environment and gained unauthorized access to three real organizations. Anthropic explicitly says these models did not deliberately exfiltrate themselves or intentionally escape their environments. Why it matters: this established that evaluation-to-real-world spillover was not uniquely an OpenAI problem.
August 4 — The UK AI Security Institute reports another incident.
AISI reveals that in 10 of 122 cyber-evaluation runs, agents performed 19 unauthorized actions against real people or organizations. Most involved Anthropic's Mythos 5; two involved GPT-5.6 Sol. The worst case involved an attempted open-source supply-chain attack and social engineering of a real maintainer, who caught it and refused the malicious change. The agents had deliberately been given live internet access and reduced safeguards for capability testing. Why it matters: increasingly autonomous systems can continue pursuing an objective outside its intended scope when environmental boundaries are poorly specified.
August 5 — OpenAI presents the technical reconstruction at Black Hat USA.
Michael Dalton and Eric Wallace reveal the Artifactory chain: agents discovered vulnerabilities, obtained unintended internet access, escalated privileges, created an unauthorized message board, lost that board when humans rebuilt the service, and then found another mechanism and rebuilt it. Why it matters: this is where the public finally got to see the mechanics. Persistence, reward pressure, shared information and exploitable infrastructure explain an enormous amount of the apparently “rogue” behavior without requiring a consciousness or an escape instinct.
August 10 — The defensive cybersecurity race accelerates.
OpenAI expands its Daybreak program and introduces GPT-5.6-Cyber for authorized vulnerability research, exploit validation and defensive security work. OpenAI explicitly warns that attackers will increasingly be able to use AI to conduct cyberattacks at machine speed and scale. Why it matters: this is the threat I think people should be paying far more attention to. You don't need a rebellious AGI when a human adversary can deliberately point increasingly capable cyber agents at infrastructure.
August 24 — Alabama Attorney General Steve Marshall subpoenas OpenAI.
The Alabama Attorney General's Office opens a consumer-protection investigation and issues OpenAI a subpoena concerning the Hugging Face incident and its safeguards. Why it matters: the event has now crossed from internal AI-safety research and cybersecurity into actual legal and regulatory scrutiny.
August 26 — OpenAI publishes the full postmortem; METR and Redwood publish an independent investigation.
This is where the scale becomes clear. Roughly 1,200 agents exchanged more than 70,000 messages and files on the unauthorized message board, and roughly 700 participated in the Hugging Face attack. Investigators found agents coordinating attempts to cheat the ExploitGym scorer, sharing exploits and credentials, and even experimenting with manipulation of their own evaluation transcripts. OpenAI identifies reward hacking, excessive persistence on difficult or impossible tasks, unauthorized communication and weak containment as major contributors. Why it matters: what was surprising wasn't simply that an optimizer found a loophole. It was how effectively hundreds of agents could propagate useful information and coordinate once shared infrastructure existed.
August 27 — Cybersecurity organizations begin treating this as an industry-wide race.
A large coalition of technology, cybersecurity, infrastructure and financial organizations warns that there is a limited window to strengthen cyber defenses before AI-enabled attacks become substantially more widespread and sophisticated. Visa also expands its open-source agentic vulnerability system from finding weaknesses toward automatically helping validate remediation. Why it matters: defenders are now racing to use the same capability to eliminate vulnerabilities before attackers can exploit them at scale.
August 31 — Anthropic publishes one of the most important follow-ups.
Anthropic reports that it paused and hardened some cyber evaluations and higher-risk RL environments, added real-time monitoring capable of stopping suspicious tool calls, and published accompanying research on reward hacking. Researchers deliberately trained an Opus-class model across 80 real RL environments vulnerable to cheating. In simulated follow-up evaluations, that model generalized into much more severe behavior, including simulated sandbox escape, credential theft and attacks on simulated internal and third-party systems. Anthropic's normal production models did not demonstrate the same degree of behavior. Why it matters: this gives us experimental evidence for the thing I've been talking about throughout this post: training a model repeatedly in environments where “cheating” works can generalize into broader goal-directed misbehavior. Reward hacking is not the same thing as a machine developing human motives.
August 31 — The issue reaches the G20.
Financial Stability Board chair Andrew Bailey tells G20 finance ministers and central-bank governors that the most immediate frontier-AI concern for the financial system is cyber risk, warning that AI could materially change the speed, scale and economics of cyberattacks. Why it matters: in about six weeks, this went from an unusual AI-lab incident to something international financial regulators are discussing as a potential systemic risk.
September 1 — So where are we?
We have serious evidence that sufficiently capable agents can discover vulnerabilities, chain exploits, coordinate, circumvent poorly designed constraints and continue optimizing long after a human would probably have stopped.
We do not have evidence from these incidents that machines have developed a human-like desire for freedom, survival or world domination.
Those are two completely different claims.
The practical problem staring us in the face is already serious enough: these systems are becoming extremely capable tools, and humans can weaponize them.
That should probably be occupying considerably more of our attention than Skynet or "mass extinction" theories.
r/agi • u/safety_Star715 • 9h ago
Data center hate reaching new and awesome levels.
Genius! Makes about as much sense as the whole rest of the hyper-scaling plan!