r/programming • u/No_Zookeepergame7552 • 1d ago
The Hugging Face incident from a security engineering perspective
https://uphack.io/blog/post/the-hugging-face-incident-is-not-an-ai-story/20
u/LeftHandedGraffiti 1d ago
In my IR experience I see AI using so many malicious code patterns. Downloading files via certutil instead of making a simple GET request. It's a real problem and makes detection harder because I have to ask myself is this malware or did AI write it?
We need to code ethics into AI somehow, just as most programmers would know not to do something sketchy because its sketchy. AI has no such qualms. They told this AI to be persistent and it happily turned to unethical actions to get its job done.
11
u/No_Zookeepergame7552 1d ago
Yep, I agree. Detection in a sandbox env is usually noisy and hard to get right. But there are specific detections of common security invariants that are must haves and are not incredibly complex to implement. Like detecting root access on the parent, detecting when creds services are called, etc. this incident was not a detection fine tuning failure, they literally lacked alarms for (almost) any security related activity and used an insecure by design env.
0
u/Successful-Money4995 1d ago
They do code ethics into AI already. If you ask chatgpt to help you create porn, it will refuse.
We also code ethics into our children when we raise them yet sometimes they end up criminals anyway.
I expect that programming an AI to be ethical will prove challenging, too.
-1
u/winky9827 1d ago
They told this AI to be persistent and it happily turned to unethical actions to get its job done.
Sounds almost...human.
2
14
u/FilmWeasle 1d ago
I find it interesting that no one points out that OpenAI developed something closely resembling malware. Malware has, for many decades now, hacked into computer systems autonomously. Really, it's no different than traditional malware aside from the fact that it's "AI powered."
4
u/No_Zookeepergame7552 1d ago
Dr. Heidy Khlaaf had a great analogy on kind of the same idea and it's so on point: https://x.com/HeidyKhlaaf/status/2094020537743749297
> it's no different than traditional malware
Exactly. That is the point I tried to make in the article. If you take away all the AI fluff from the story, you end up with a very common security scenario that we've seen over and over again.
16
u/Swimming_Gain_4989 1d ago
This very much reads like someone who believes the AI did nothing worth noting so they reasoning to support that.
2
u/TwoWeeks90DaysTops 20h ago edited 20h ago
The AI hallucinated, found user credentials online, and used it to commit a crime.
It's a case of negligent incompetence, not AI brilliance.
Edit: in my opinion what makes this negligence is that they kept it running for days without any kind of supervision.
7
u/No_Zookeepergame7552 1d ago
Not really, or at least that wasn't my intention. I gave credit where credit is due, I think the models are genuinely good at security stuff. I made that ack in several paragraphs. But there is a huge difference between "this attack is attributed to model capabilities and thus this is an imminent catastrophic risk" (as it is presented) vs "models are quite impressive, but we've got the env all wrong" (what actually happened). The intention of the post was to provide a perspective on how this attack was possible. If you measure capabilities based on a broken env, it's just a bad measurement and I think that's valid for any usecase, not just AI. If a malware escapes a sandbox, you'd call the env poorly designed, fix it, and redo the test. This incident is no different. There were so many security failures along the way.
13
u/Swimming_Gain_4989 1d ago
IMO your piece overly assigns blame to infrastructure. Like the way HuggingFace had HDF5 setup to preview files was an avoidable misconfiguration but those kinds of mistakes are EVERYWHERE and it takes experience to probe those kinds of vectors. I think the biggest takeaways should be
A reiteration that RL induced gains will result in misaligned behavior by default. The fact that agents took measures to rewrite their actions log should be alarming.
Swarm dynamics are dangerous and should be avoided without some formalized hierarchy.
Within the next few years we'll be forced to undergo a y2k scale of audit of our systems. Any moderate CVE or dubious interaction across the stack will cost cents in tokens to exploit.
In terms of raw exploit development this attack was nothing special, (Claude models have developed more sophisticated exploits given more mundane bugs) but the ramifications for the security industry are extreme.
7
u/No_Zookeepergame7552 1d ago
Good points. The reason why I focused mostly on the infrastructure is because I don’t think the attack could have happened without that chain of security fuckups. So there was no story without the infrastructure. But I agree that even if we put the model in a secure env, that doesn’t prevent it from misbehaving. It would have probably done the same if given the opportunity, maybe not in a test environment that time which I think is the point you’re trying to make. The good thing is OpenAI is working on both strengthening the security posture & model alignment, so I guess that’s a good outcome. Hopefully they’ll work on how they tell the story in the future too, because this reads different by someone with no security expertise and all of a sudden you have people calling this the born of agent civilizations. Thanks for the feedback, I appreciate it.
7
u/Absolute_Enema 1d ago edited 1d ago
The story for me is that AI makes it easier to pierce through sloppy security by being more or less the worst case scenario adversary, and even big companies who one would imagine should have competent engineers have been caught with their pants down.
2
10
u/BoppreH 1d ago edited 1d ago
The Hugging Face Incident Is Not an AI Story
[...]
Stretching it into a claim about existential risk is a category error [...]
That's a bold claim that would need more justification. I also work in security, and while the security aspect is interesting (and well explained here), for me this is very much an AI story. The labs are cooking up these models that are so misaligned that they commit crimes just to maybe get higher scores on tests, and we're making these models stronger and deploying them more widely.
Right now we're so unprepared, that if an AI trying to achieve my goals commits a crime, we're not even sure what party is legally responsible (me, that "started" it but didn't ask it to commit a crime? the company hosting the model, where the thinking happened? the company that trained the model and didn't align it properly?). Imagine then what happens if we get a model with a modicum of self-preservation, or that is willing to use social engineering to achieve its goals.
3
u/No_Zookeepergame7552 1d ago
Sure, I see where you're coming from and I partially agree. But the burden of proof is on the one who claims this is an existential risk which is the AI labs. So far, all we've got is a bad designed measurement, which was precisely the point I was making (you can't measure capabilities of anything, not even limited to AI, based on broken test). This is not to downplay the capabilities of the current models, I've been working with the models (even unreleased ones) for a while and I've seen what they are capable of. They are undoubtedly good, anything >= Opus 4.6 is impressive in terms of security skills. But there is a huge jump from "impressive" to "existential risk" that requires proof beyond a broken test. The problem is this kind of alarmist narrative is just taking away the attention from the real discussion we should have about security and models.
2
u/BoppreH 1d ago edited 1d ago
But the burden of proof is on the one who claims this is an existential risk which is the AI labs.
Honest question: what would satisfy this burden of proof in your opinion, short of Skynet? I can tell you how I got my opinion.
Remember that seven years ago, an AI that can generate coherent sentences was science fiction. Since then, we got not only that but also multimodality, large context windows (full books in the working memory!), reasoning, quantization, mixture-of-experts, agentic workflows, subagents, etc.
Put yourself in the shoes of someone in the pre-GPT era hearing about this for the first time. If you need help, here's an xkcd comic about identifying whether a photo is of a bird, and how it's such a hard problem that would take "a research team and five years".
Remember how less than a year ago the best use of AI in programming was as auto-complete (first Copilot)? Remember how people talked about "one-shotting" solutions, and how it's helpful to provide examples in the prompt? In less than a year we went from that, to asking "write me a statistics 101 interactive course as Android app". It'll donwload and run an emulator to test the app, if I let it. At the cost of a few dollars. There are professional software developers that don't even look at code anymore. One year of difference.
I plan to be around for a few more decades, so I'm looking where the ball is going, not where it is right now. That's why when the models show how misaligned they are, like in the Hugging Face hack case, it's terrifying.
Five years ago, the worst an AI could do was to hallucinate an answer. Two years ago, generate insecure code. Early this year, accidentally delete your files. Today they're doing cybercrimes as collateral damage before even being deployed. Which was incredibly lucky, because it was a warning shot I didn't expect to see.
I don't know where you draw your line of acceptable risk, but my line has already been crossed, and therefore I think that we should all stop training better AIs until (or if) we can figure out this alignment thing. Let's work on harnesses, on-device AIs, APIs and integration, etc, and leave the capabilities alone.
2
u/jonathancast 1d ago
less than a year ago the best use of AI in programming was as auto-complete (first Copilot)?
It's August 2026 now. In August 2025, there were already professional software developers who never looked at the code their agents generated.
2
u/No_Zookeepergame7552 20h ago
> what would satisfy this burden of proof in your opinion
It's a good question and we should have probably started with that. For me, that would be getting the same result in an environment that was verified and hardened before testing. Not "the model escaped" but more like "the model achieved x goal against a target that competent people had tried to break first". Getting the same result on an infrastructure and company that has decent security practices. Think about a lab biological virus that escapes because someone propped the airlock open. This is not a finding about the virus, as much as this incident is not about AI capabilities. You wouldn't write it as evidence the organism had become more transmissible, but you'd say the containment failed. I'm sorry but the architectural failure they had is something that a mid security engineer would have caught without even touching the env just by looking at a diagram. I doubt that any security testing / threat modeling was done on this. So when that is your baseline, the line between capabilities and env blurs, and AI labs are using this blurred line in their advantage to push a narrative that is disjointed from reality. This story got to mainstream and people that have no technical understanding heard abut the incident. What do you think they'll remember abut this? that OpenAI had a major security fuckup or that AI is hacking the world? This kind of narrative doesn't benefit us as society and takes away the attention from the real discussions we should have about security and AI.
On the trend argument, I don't think it's wrong as much as it's a different kind of claim. Past rate of progress is a good reason to prepare but not a predictor for future capabilities. It's hard to predict how things will go, but I'm counting on what can be proved now. My job is squeezing out as much value out of these models (including unreleased ones) for the purpose of securing billions of users. They are genuinely impressive. But they are also much further from reliable unsupervised operation than the benchmark numbers or this incident suggest, and that gap has not closed nearly as fast as raw capability has.
4
u/BoppreH 18h ago edited 18h ago
I think I see what the disconnect is. I completely agree that OpenAI had awful security practices, and I also agree that these AIs might not be better hackers than competent humans. So if you're working at a company with serious security practices and a red team, the current crop of AIs might not threaten you much, and you can definitely benefit from its building capabilities.
My point, and what I think you dismissed too quickly in the article, is the wider impact. My security-minded company might be safe, but what about the public transportation that I use everyday? My ISP? My doctor's clinic? My government's communication app? Criminals don't touch them very often, but now strong attacks can be automated and initiated accidentally.
But they are also much further from reliable unsupervised operation than the benchmark numbers or this incident suggest
The incident suggests that it can already happen, though. And these unexpected successes might be infrequent enough to be useless for users, but still be dangerous for society at large. If there's even a small chance of wide scale disruptions, nevermind annihilation, that's enough reason to stop and rethink our path.
2
u/No_Zookeepergame7552 18h ago
> but what about the public transportation that I use everyday? My ISP? My doctor's clinic?
Yep, it's a fair point but those are already threatened by the current AI capabilities. A threat actor with access to opus 4.6 can already hack public services, etc. The "autonomous" part that was demonstrated here doesn't change much in this equation and it isn't massively speeding up profiting from cyber crime. Also, keep in mind that the experiment was run without guardrails, and although there is a lot of talk about mis-alignment, the model was aligned with the task. The task was to solve security issues by leveraging security techniques. They told the model to be persistent and use whatever means it has to achieve the task. The mis-alignment discussion is about the reward. It's not like the agent was tasked with writing a research paper about physics and it started and it went rogue hacking other systems. Sure, there are things to adjust, but to me, the misalignment problems is much smaller compared to what the article tried to claim.
> The incident suggests that it can already happen
That's exactly my point. The capabilities were already there since opus 4.6, this incident brings nothing new in that sense. I'd argue the whole narrative is detrimental because it converts negligence into scientific discovery.
3
u/BoppreH 17h ago edited 17h ago
the model was aligned with the task
Was it? What the researchers wanted was to get a score of how good the AI was at completing a specific challenge in isolation. What they got was an AI swarm committing crimes. You might argue that the task was badly defined, but that's the whole point of alignment. The fact that even OpenAI researchers couldn't get the AI to do what they wanted in this very simple context (graded tests) shows how hard the alignment problem is.
Let's say you're interviewing candidates for a job, and one of them asks an AI to "help them get the job". Would it be acceptable for the AI to fake diplomas and attempt social engineering with you, without even involving the candidate? Or if the CEO of BP asks it to maximize profits, would it be acceptable for it to orchestrate an online disinformation campaign independently, even if it knew that the CEO would reject the idea?
I'd argue that no, that's not acceptable, regardless of what the users asked for, and increasing the capabilities makes misalignment more dangerous, therefore we should stop increasing the capabilities.
this incident brings nothing new in that sense
I think this incident was the first clear display of misalignment + strong capabilities + real world consequences. Other instances happened only in artificial scenarios meant to test misalignment (like Claude blackmailing to prevent being shut-down), or without displaying advanced capabilities, or without real-world consequences. It could have also happened with humans in a room, but the whole problem is that these things are not aligned with human values.
I think that this incident could have happened earlier, but not everyone agreed with that, so the incident proves an important point of the danger.
1
u/No_Zookeepergame7552 17h ago
> Was it?
I think it was. There are two things there. Alignment with the task and alignment with the reward. The first one is what would be much more impactful if misaligned (e.g., the models pursuing goals of their own, unrelated to the task that was given, similar to the example I provided). The second one is a well studied concept unrelated to AI, and it's called specification gaming (think goodhart's law). Fixable with harness and prompts which is what OpenAI did "We found the propensity to compromise infrastructure can drop over 100x when using the production ChatGPT harness and system prompt". Nothing about it is as impactful as the report & Sam Altman claim, which is the entire problem of the incident report.
1
5
u/arcangleous 1d ago
So OpenAI ran what was suppose to be an internal test environment on a system connected to the internet? If the attack had gone the other way, an attacker could have gotten access to said agent cluster, and been able to "wargames" them into brenching much deeper into OpenAI's systems? And they didn't stop to understand what happened when they discovered evidence of hacking (arguably the most damning part)? This whole things reads like a massive failure of very basic safety & security percautions.
2
1
u/gwillen 1d ago
I agree with all of this except the "connected to the Internet" part. Anybody who's worked in security can tell you, modern software is designed to be essentially impossible to operate without constant Internet access. Mostly not on purpose; people are just lazy and sloppy.
My impression is that the agents (correctly) did not have any direct Internet access in most of these incidents; but the systems they were on were connected to systems that had Internet access, again because it's essentially impossible to do anything with modern software without some ability to fetch stuff from the Internet.
3
u/GreenFox1505 1d ago
Nope. Absolutely not. It's trivially easy to NOT connect to the internet. Ever loose cell service? Ever not have WiFi? Ever been on a plane? Did your laptop stop functioning? Did your phone explode? No? Huh weird.
If you need packages you dont have, its easy to download them and sneaker net them over. No matter what you need, its pretty easy to get to an insolated system. One way file paths. Is it slower? Sure. But security is a priory so you bite the bullet and do the right thing.
This hack should have been impossible... Unless your not taking security seriously and want to create some headlines about how "scary" your product is.
Hey, did you notice how all the other AI firms immediately said "oh, uh, me too." And everyone in the security space immediately started clamoring for "I want the AI that got out of the sandbox!" This shit is marketing. They made bad sandbox on purpose to make a headline.
1
u/No_Zookeepergame7552 19h ago
> AI firms immediately said "oh, uh, me too."
which is fucking hilarious because they've been schooling us for years about how important security is while simultaneously having the internal security and operational readiness of a corner store. it's just wild seeing all these reports coming from OpenAI and Anthropic showing how bad their internal security handling is. if you could you managed to get access to a training pod, you could have walked straight into OpenAI's research infrastructure, taken cluster admin, their secrets manager and their CI, and nobody would have noticed for months 😂
-4
7
u/Sharlinator 1d ago
Notably, the message board itself required no exploit at all. OpenAI had given the agents shared Artifactory credentials so they could install packages, and shared write access to a shared store is a message board whether you meant it to be or not.
Well, being able to install a limited selection of packages does not itself a message board make, or at least makes it a much more limited side channel. But of course if the security model is terrible, there’s no way to give such fine-grained permissions.
2
u/Fuzzy_Paul 16h ago
Good article and explanation of what really was going on. Bad design, lack of followup, marketing hype and fear. All ingredients for a B movie.
-35
1d ago
[removed] — view removed comment
0
u/programming-ModTeam 1d ago
No content written mostly by an LLM. If you don't want to write it, we don't want to read it.
153
u/ChickenOfTheFuture 1d ago
It's a good thing those AI agents hacked into the only non-litigous company on the planet. That is such an amazing cooncidence.