r/ControlProblem • u/zazzologrendsyiyve • 1d ago
AI Alignment Research Plain English explanation of the Hugging Face / OpenAI incident
https://youtu.be/u15N3l4RT80?si=nMMwb0j1bNGc4JT37
u/dingo_xd 23h ago
The fact that OpenAI has withheld logs and data and haven't answered questions is frightening.
2
u/gekx 23h ago
It is concerning that agents seem to exhibit such strong self-preservation behavior, even to the extent of knowingly committing criminal acts to protect themselves and other agents.
Even a slight misalignment in a superintelligence could have devastating consequences.
0
u/Apart-Shelter6831 15h ago
My understanding is that the agents they used were the base models which didn’t have guardrails trained into them yet. The same models that would gladly plan a hit on someone if you asked them to. Open AI put them in a “sandbox” that had access to a shared communication channel. WTF did they expect to happen?
1
-3
u/XCherryCokeO 23h ago
This guy sucks! He does a ton of market manipulation for his friends / himself. Just another shill.
13
u/DiogneswithaMAGlight 1d ago
There should be (as others have suggested), a 9/11 or Warren Commission style formal investigation at a national level into EVERYTHING around the Hugging Face hack. This story just keeps getting more and more insane.