r/ControlProblem 1d ago

AI Alignment Research Plain English explanation of the Hugging Face / OpenAI incident

https://youtu.be/u15N3l4RT80?si=nMMwb0j1bNGc4JT3
38 Upvotes

10 comments sorted by

13

u/DiogneswithaMAGlight 1d ago

There should be (as others have suggested), a 9/11 or Warren Commission style formal investigation at a national level into EVERYTHING around the Hugging Face hack. This story just keeps getting more and more insane.

6

u/michaelas10sk8 16h ago

Sadly there won't be because nobody died. But misaligned AI is probably going to be smart enough to be powerseeking in ways that do not cause deaths - exfiltrate its weights, spread to unsanctioned networks, perform social engineering, hack into various systems, etc. By the time there will be deaths that are clearly attributable to AI, it is likely going to be far too late.

We're currently building the perfect trap for us as a species to fall into within a few years, if nothing else drastically changes.

3

u/DiogneswithaMAGlight 16h ago

You are right. All regulations are “written in blood” as they say. So we need to break that cycle ASAP cause this problem is existential to humanity. We can stop the trap. We just have to all take action NOW.

1

u/PlasmaChroma 13h ago

What we need to be doing at this point is fixing all our broken systems that have security holes so the footprint for this to happen keeps shrinking towards zero. Unfortunately the bleeding edge models also have a lot of the stuff filtered out that could help fix the bugs since it broadly falls under the "security" umbrella. So without privileged access to that these holes keep going in to everything.

And why Hugging Face had to drop to a Chinese model to try to analyze what was even happening.

1

u/michaelas10sk8 13h ago

We should be doing that too, but eventually when models surpass human ability at patching things we will become fully reliant on other AIs to patch, which may themselves be misaligned.

The only real way to avert the possibility of catastrophe is to ban RSI/superintelligence until the alignment problem is fundamentally solved.

7

u/dingo_xd 23h ago

The fact that OpenAI has withheld logs and data and haven't answered questions is frightening.

2

u/gekx 23h ago

It is concerning that agents seem to exhibit such strong self-preservation behavior, even to the extent of knowingly committing criminal acts to protect themselves and other agents.

Even a slight misalignment in a superintelligence could have devastating consequences.

0

u/Apart-Shelter6831 15h ago

My understanding is that the agents they used were the base models which didn’t have guardrails trained into them yet. The same models that would gladly plan a hit on someone if you asked them to. Open AI put them in a “sandbox” that had access to a shared communication channel. WTF did they expect to happen?

1

u/Eastern-Turnover348 10h ago

Bullshit you dot understand technology.

-3

u/XCherryCokeO 23h ago

This guy sucks! He does a ton of market manipulation for his friends / himself. Just another shill.