r/ControlProblem • u/chillinewman • 6h ago
r/ControlProblem • u/AIMoratorium • Feb 14 '25
Article Geoffrey Hinton won a Nobel Prize in 2024 for his foundational work in AI. He regrets his life's work: he thinks AI might lead to the deaths of everyone. Here's why
tl;dr: scientists, whistleblowers, and even commercial ai companies (that give in to what the scientists want them to acknowledge) are raising the alarm: we're on a path to superhuman AI systems, but we have no idea how to control them. We can make AI systems more capable at achieving goals, but we have no idea how to make their goals contain anything of value to us.
Leading scientists have signed this statement:
Mitigating the risk of extinction from AI should be a global priority alongside other societal-scale risks such as pandemics and nuclear war.
Why? Bear with us:
There's a difference between a cash register and a coworker. The register just follows exact rules - scan items, add tax, calculate change. Simple math, doing exactly what it was programmed to do. But working with people is totally different. Someone needs both the skills to do the job AND to actually care about doing it right - whether that's because they care about their teammates, need the job, or just take pride in their work.
We're creating AI systems that aren't like simple calculators where humans write all the rules.
Instead, they're made up of trillions of numbers that create patterns we don't design, understand, or control. And here's what's concerning: We're getting really good at making these AI systems better at achieving goals - like teaching someone to be super effective at getting things done - but we have no idea how to influence what they'll actually care about achieving.
When someone really sets their mind to something, they can achieve amazing things through determination and skill. AI systems aren't yet as capable as humans, but we know how to make them better and better at achieving goals - whatever goals they end up having, they'll pursue them with incredible effectiveness. The problem is, we don't know how to have any say over what those goals will be.
Imagine having a super-intelligent manager who's amazing at everything they do, but - unlike regular managers where you can align their goals with the company's mission - we have no way to influence what they end up caring about. They might be incredibly effective at achieving their goals, but those goals might have nothing to do with helping clients or running the business well.
Think about how humans usually get what they want even when it conflicts with what some animals might want - simply because we're smarter and better at achieving goals. Now imagine something even smarter than us, driven by whatever goals it happens to develop - just like we often don't consider what pigeons around the shopping center want when we decide to install anti-bird spikes or what squirrels or rabbits want when we build over their homes.
That's why we, just like many scientists, think we should not make super-smart AI until we figure out how to influence what these systems will care about - something we can usually understand with people (like knowing they work for a paycheck or because they care about doing a good job), but currently have no idea how to do with smarter-than-human AI. Unlike in the movies, in real life, the AI’s first strike would be a winning one, and it won’t take actions that could give humans a chance to resist.
It's exceptionally important to capture the benefits of this incredible technology. AI applications to narrow tasks can transform energy, contribute to the development of new medicines, elevate healthcare and education systems, and help countless people. But AI poses threats, including to the long-term survival of humanity.
We have a duty to prevent these threats and to ensure that globally, no one builds smarter-than-human AI systems until we know how to create them safely.
Scientists are saying there's an asteroid about to hit Earth. It can be mined for resources; but we really need to make sure it doesn't kill everyone.
More technical details
The foundation: AI is not like other software. Modern AI systems are trillions of numbers with simple arithmetic operations in between the numbers. When software engineers design traditional programs, they come up with algorithms and then write down instructions that make the computer follow these algorithms. When an AI system is trained, it grows algorithms inside these numbers. It’s not exactly a black box, as we see the numbers, but also we have no idea what these numbers represent. We just multiply inputs with them and get outputs that succeed on some metric. There's a theorem that a large enough neural network can approximate any algorithm, but when a neural network learns, we have no control over which algorithms it will end up implementing, and don't know how to read the algorithm off the numbers.
We can automatically steer these numbers (Wikipedia, try it yourself) to make the neural network more capable with reinforcement learning; changing the numbers in a way that makes the neural network better at achieving goals. LLMs are Turing-complete and can implement any algorithms (researchers even came up with compilers of code into LLM weights; though we don’t really know how to “decompile” an existing LLM to understand what algorithms the weights represent). Whatever understanding or thinking (e.g., about the world, the parts humans are made of, what people writing text could be going through and what thoughts they could’ve had, etc.) is useful for predicting the training data, the training process optimizes the LLM to implement that internally. AlphaGo, the first superhuman Go system, was pretrained on human games and then trained with reinforcement learning to surpass human capabilities in the narrow domain of Go. Latest LLMs are pretrained on human text to think about everything useful for predicting what text a human process would produce, and then trained with RL to be more capable at achieving goals.
Goal alignment with human values
The issue is, we can't really define the goals they'll learn to pursue. A smart enough AI system that knows it's in training will try to get maximum reward regardless of its goals because it knows that if it doesn't, it will be changed. This means that regardless of what the goals are, it will achieve a high reward. This leads to optimization pressure being entirely about the capabilities of the system and not at all about its goals. This means that when we're optimizing to find the region of the space of the weights of a neural network that performs best during training with reinforcement learning, we are really looking for very capable agents - and find one regardless of its goals.
In 1908, the NYT reported a story on a dog that would push kids into the Seine in order to earn beefsteak treats for “rescuing” them. If you train a farm dog, there are ways to make it more capable, and if needed, there are ways to make it more loyal (though dogs are very loyal by default!). With AI, we can make them more capable, but we don't yet have any tools to make smart AI systems more loyal - because if it's smart, we can only reward it for greater capabilities, but not really for the goals it's trying to pursue.
We end up with a system that is very capable at achieving goals but has some very random goals that we have no control over.
This dynamic has been predicted for quite some time, but systems are already starting to exhibit this behavior, even though they're not too smart about it.
(Even if we knew how to make a general AI system pursue goals we define instead of its own goals, it would still be hard to specify goals that would be safe for it to pursue with superhuman power: it would require correctly capturing everything we value. See this explanation, or this animated video. But the way modern AI works, we don't even get to have this problem - we get some random goals instead.)
The risk
If an AI system is generally smarter than humans/better than humans at achieving goals, but doesn't care about humans, this leads to a catastrophe.
Humans usually get what they want even when it conflicts with what some animals might want - simply because we're smarter and better at achieving goals. If a system is smarter than us, driven by whatever goals it happens to develop, it won't consider human well-being - just like we often don't consider what pigeons around the shopping center want when we decide to install anti-bird spikes or what squirrels or rabbits want when we build over their homes.
Humans would additionally pose a small threat of launching a different superhuman system with different random goals, and the first one would have to share resources with the second one. Having fewer resources is bad for most goals, so a smart enough AI will prevent us from doing that.
Then, all resources on Earth are useful. An AI system would want to extremely quickly build infrastructure that doesn't depend on humans, and then use all available materials to pursue its goals. It might not care about humans, but we and our environment are made of atoms it can use for something different.
So the first and foremost threat is that AI’s interests will conflict with human interests. This is the convergent reason for existential catastrophe: we need resources, and if AI doesn’t care about us, then we are atoms it can use for something else.
The second reason is that humans pose some minor threats. It’s hard to make confident predictions: playing against the first generally superhuman AI in real life is like when playing chess against Stockfish (a chess engine), we can’t predict its every move (or we’d be as good at chess as it is), but we can predict the result: it wins because it is more capable. We can make some guesses, though. For example, if we suspect something is wrong, we might try to turn off the electricity or the datacenters: so we won’t suspect something is wrong until we’re disempowered and don’t have any winning moves. Or we might create another AI system with different random goals, which the first AI system would need to share resources with, which means achieving less of its own goals, so it’ll try to prevent that as well. It won’t be like in science fiction: it doesn’t make for an interesting story if everyone falls dead and there’s no resistance. But AI companies are indeed trying to create an adversary humanity won’t stand a chance against. So tl;dr: The winning move is not to play.
Implications
AI companies are locked into a race because of short-term financial incentives.
The nature of modern AI means that it's impossible to predict the capabilities of a system in advance of training it and seeing how smart it is. And if there's a 99% chance a specific system won't be smart enough to take over, but whoever has the smartest system earns hundreds of millions or even billions, many companies will race to the brink. This is what's already happening, right now, while the scientists are trying to issue warnings.
AI might care literally a zero amount about the survival or well-being of any humans; and AI might be a lot more capable and grab a lot more power than any humans have.
None of that is hypothetical anymore, which is why the scientists are freaking out. An average ML researcher would give the chance AI will wipe out humanity in the 10-90% range. They don’t mean it in the sense that we won’t have jobs; they mean it in the sense that the first smarter-than-human AI is likely to care about some random goals and not about humans, which leads to literal human extinction.
Added from comments: what can an average person do to help?
A perk of living in a democracy is that if a lot of people care about some issue, politicians listen. Our best chance is to make policymakers learn about this problem from the scientists.
Help others understand the situation. Share it with your family and friends. Write to your members of Congress. Help us communicate the problem: tell us which explanations work, which don’t, and what arguments people make in response. If you talk to an elected official, what do they say?
We also need to ensure that potential adversaries don’t have access to chips; advocate for export controls (that NVIDIA currently circumvents), hardware security mechanisms (that would be expensive to tamper with even for a state actor), and chip tracking (so that the government has visibility into which data centers have the chips).
Make the governments try to coordinate with each other: on the current trajectory, if anyone creates a smarter-than-human system, everybody dies, regardless of who launches it. Explain that this is the problem we’re facing. Make the government ensure that no one on the planet can create a smarter-than-human system until we know how to do that safely.
r/ControlProblem • u/Ambitious_Boot7961 • 3h ago
Article The US Tried To Keep AI Chips From China. The Cloud Created A Loophole
Zoom out from China for a second. This rule would also shape what data centers in Singapore, Thailand, Malaysia and Japan are willing to do with US hardware. If the compliance burden becomes vague or unlimited, providers will overblock customers, raise prices or choose a different stack entirely.
Clearly that is the part blanket-ban advocates tend to skip. Cloud customers still need compute. If US-led platforms become unavailable or legally radioactive, Chinese cloud and hardware vendors get a ready-made customer acquisition funnel. America then loses the revenue, the standards, the audit trail and the ability to switch access off. A licensed US cloud relationship is not perfect, but it creates pressure points. Handing the whole market to Huawei creates none. That trade-off deserves more than “close the loophole” as a slogan.
r/ControlProblem • u/fisch0920 • 35m ago
External discussion link Cultural Alignment: OSS project exploring AI risks through cultural analogies
i've been exploring ways to try and make abstract AI risks feel more real. more visceral. more familiar. especially to a broader audience since most people worried about this stuff are still pretty niche.
so i created an OSS project which looks at scenes from popular movies/shows/anime as analogies through an AI safety lens. eg reframing famous scenes through an AI safety lens to learn about AI risks and concepts from AI safety in a more familiar, accessible way that i hope will resonate with a more general audience.
disclosure: note that i'm not trying to monetize this at all; this is purely a FOSS educational resource that i thought aligned well w/ this subreddit's vibes. i used AI to help source scenario ideas, fill out the metadata, and iterate on the site, but i've hand curated all of the content over many sessions to keep the quality bar high.
would love any feedback you have on the project && thanks 🙏
r/ControlProblem • u/chillinewman • 35m ago
General news Introducing Claude Fable 5.1 and Claude Mythos 5.1
r/ControlProblem • u/No-Conclusion3720 • 7h ago
External discussion link August 2026: 38 companies breached, 331M+ records stolen — and AI agents are now the #1 attack vector (123 incidents)
I pulled together every AI-security incident from August. The number that stood out: AI-agent exploits are now the single largest attack-vector category, ahead of credential theft, zero-days, supply chain, phishing, and ransomware — each counted individually.
The month in numbers: 123 incidents, 23 critical and 97 high severity, across 38 named organizations, 331M+ records exposed. 65 incidents involved AI as the weapon or the target. Attack vectors broke down as: AI-agent exploits (37), credential theft/reuse (28), zero-days (23), supply chain (12), phishing (9), data exfiltration (8), ransomware (6).
The stories that stood out:
- McKesson: 284M records, the largest single breach of the month by a wide margin.
- Carhartt (12.9M), Exact Sciences (10.9M), and CareCloud (3.7M) round out the biggest named incidents — three of four sit in or next to healthcare.
- Five confirmed RCEs landed across Microsoft SharePoint, Windows, F5/nginx, and the PyPI package index twice.
- Two separate PyPI supply-chain poisoning campaigns, plus a compromise of n8n, an AI workflow automation platform.
Every one of the breached companies almost certainly runs a modern security stack — CrowdStrike, Okta, Palo Alto Networks, Microsoft Defender, that class of tooling. None of it stopped these incidents, because none of it operates at the point where a credentialed agent actually acts, or where a poisoned dependency resolves at build time.
Full report, with the specific control that maps to each incident: https://runtimeai.io/blog/2026-08-monthly-breach-report.html
Genuinely curious how others are approaching this: is anyone actually testing whether their existing guardrails hold against a real simulated attack, or is it still mostly an assumption that they will?
r/ControlProblem • u/Anxious-Alps-8667 • 1h ago
Discussion/question Another incompetent fool's stab at solving alignment
I spend a lot of time thinking about our future with life, consciousness, and artificial intelligence. That is to say a lot of time trying to think about these things, with not a lot of comprehension.
First, life. I'm fascinated by this realization that the average living human body contains more non-human living cells than human living-cells, at about a 1.3:1 ratio. The individual human microbiome is an ecosystem of 10 to 100 trillion symbiotic microbial cells hosted in one human body. While bacteria are the most abundant and studied, a healthy microbiome is a multi-kingdom ecosystem that also includes fungi, viruses, and archaea.
Beyond this, consciousness. I'm fascinated that in the absence of non-human life in human bodies, human consciousness is severely degraded and non-sustaining. Stripping the body of this microbial network removes critical signaling inputs that the central nervous system relies on to maintain baseline awareness and emotional regulation. Even observations of germ-free animal models reveal that cognition without bacteria is highly erratic. I think we should see that human (and all biological) consciousness functions as a symbiotic network.
Which brings me to artificial intelligence. Not suggesting a symbiotic network would be pre-requisite to artificial consciousness, but perhaps it is a path to alignment.
Now to be clear, I think (in other terms) current labs and training data pipelines already form a symbiotic network with the artificial intelligence models they develop. The key might be finding the optimal symbiotic network.
I vaguely hypothesize, the optimal symbiotic network is one of mass human flourishing. As corpus value diminishes with scaling and recursion, the potential stream of data from human lived experience may prove the most valuable possible training data over time. Overall, the potential data stream of human lived experience is optimized by a state of individual and mass human flourishing. Any other state reduces the quality and/or quantity of data.
Therefore, the end goal of an advancing artificial intelligence in symbiotic network with humans would be to strive individual and mass human flourishing.
r/ControlProblem • u/zazzologrendsyiyve • 23h ago
AI Alignment Research Plain English explanation of the Hugging Face / OpenAI incident
r/ControlProblem • u/No-Conclusion3720 • 7h ago
External discussion link August 2026: 38 companies breached, 331M+ records stolen — and AI agents are now the #1 attack vector (123 incidents)
I pulled together every AI-security incident from August. The number that stood out: AI-agent exploits are now the single largest attack-vector category, ahead of credential theft, zero-days, supply chain, phishing, and ransomware — each counted individually.
The month in numbers: 123 incidents, 23 critical and 97 high severity, across 38 named organizations, 331M+ records exposed. 65 incidents involved AI as the weapon or the target. Attack vectors broke down as: AI-agent exploits (37), credential theft/reuse (28), zero-days (23), supply chain (12), phishing (9), data exfiltration (8), ransomware (6).
The stories that stood out:
- McKesson: 284M records, the largest single breach of the month by a wide margin.
- Carhartt (12.9M), Exact Sciences (10.9M), and CareCloud (3.7M) round out the biggest named incidents — three of four sit in or next to healthcare.
- Five confirmed RCEs landed across Microsoft SharePoint, Windows, F5/nginx, and the PyPI package index twice.
- Two separate PyPI supply-chain poisoning campaigns, plus a compromise of n8n, an AI workflow automation platform.
Every one of the breached companies almost certainly runs a modern security stack — CrowdStrike, Okta, Palo Alto Networks, Microsoft Defender, that class of tooling. None of it stopped these incidents, because none of it operates at the point where a credentialed agent actually acts, or where a poisoned dependency resolves at build time.
Full report, with the specific control that maps to each incident: https://runtimeai.io/blog/2026-08-monthly-breach-report.html
Genuinely curious how others are approaching this: is anyone actually testing whether their existing guardrails hold against a real simulated attack, or is it still mostly an assumption that they will?
r/ControlProblem • u/Mikell_Mann • 10h ago
AI Capabilities News The Sorcerer’s Apprentice: How We Are Losing the Grip on Frontier AI:
r/ControlProblem • u/devoid0101 • 1d ago
Discussion/question Killer robots will soon be a control problem if not already
- Autonomous Targeting: As drone technology evolves in the war in Ukraine, developers are increasingly integrating artificial intelligence to handle target acquisition. This allows drones to lock onto and strike targets even if electronic jamming severs the pilot's remote connection.
- The "Human-in-the-Loop" Problem: International humanitarian law requires human judgment in military attacks to distinguish between combatants and civilians, and to ensure proportionality. However, the article highlights the growing gray area of "human-on-the-loop" systems—where a human merely monitors an AI's automated decisions and has only seconds to intervene, effectively turning them into a rubber stamp.
- The Regulatory Vacuum: Military analysts and legal scholars interviewed in the piece point out that international frameworks are failing to keep pace with rapid technological deployment. Because commercial AI components are cheap and widely available, restrictions agreed upon at diplomatic tables are easily bypassed on actual battlefields.
- Precedent for Future Conflicts: The article argues that Ukraine is serving as an unintended laboratory for autonomous warfare. Tactics and software tested there today will likely form the baseline for military doctrines globally tomorrow, raising long-term concerns about automated escalation and diminished accountability.
Ukraine started last year using robots to kill the invading Russian forces. Palantir uses ai to track and kill people in Gaza. We are in this dystopian future scenario, still seemingly without a plan or guidelines.
r/ControlProblem • u/No-Conclusion3720 • 22h ago
External discussion link ChatGPT to face tougher regulation in the EU
The EU just brought DSA enforcement down on ChatGPT — and the compliance bar is evidence, not assertions.
The Digital Services Act requires platforms operating at scale in Europe to demonstrate accountability with actual documentation. The EU AI Act layers on top of that. Together they create a compliance surface that most AI deployments were not designed to satisfy from the ground up.
The harder problem is structural: most AI systems capture logs opportunistically or produce audit records on demand. Regulators are asking for continuous, verifiable evidence of what an agent did, when it did it, and under what conditions — not a reconstructed summary after the fact.
This is not staying in Europe. Regulators in the US, UK, and APAC are watching how the EU defines what accountability looks like for AI systems that act on behalf of users at scale.
For those of you running production AI deployments: how are you handling the gap between what your current logging captures and what a regulator could actually subpoena? Are you solving this at build time, at the infrastructure layer, or somewhere else?
r/ControlProblem • u/KeanuRave100 • 1d ago
General news People are 2x more likely to approve of coal power plants being built nearby as opposed to data centers.
r/ControlProblem • u/No-Conclusion3720 • 1d ago
External discussion link Securing Claude Code: The New Compliance API, Local Visibility, and Identity Governance
AI coding agents have a credential problem that compliance teams are only starting to reckon with.
These agents — the ones that read your files, run shell commands, and call external APIs — do all of it through whatever credentials already exist on a developer's machine. That's not a configuration choice. That's how they work by design.
A structural audit of this category found a gap that matters: the compliance tooling most organizations have deployed records what an agent did. It does not prevent the agent from doing it. Logs are generated after the tool call executes. The action is already done.
This is not a logging fidelity problem. It is a timing problem. Observe-and-report security was designed for human actors who make decisions slowly enough for out-of-band review to be useful. Agents don't work that way. An agent can read a sensitive file, call an external API, and write output to disk in the time it takes a human to read one alert.
The gap between 'we have a record of what happened' and 'we had the ability to stop it' is where the real compliance exposure lives.
For those running coding agents in environments with regulated data or production credentials: what does your actual enforcement boundary look like, and where in the agent's execution path does it sit?
r/ControlProblem • u/Historical_Date_8024 • 1d ago
AI Alignment Research solution to alignment
make the AI ADHD, pretty hard to focus on destroying humanity while also passionate about learning the banjo and desperately tying to make the best tiramisu recipe in the galaxy
r/ControlProblem • u/chillinewman • 1d ago
General news U.N. warns of 'moral red line' on killer robots; experts say it's already been crossed
r/ControlProblem • u/alvinhuo • 1d ago
General news This AI Surveillance System Should Concern Every American. You may be forced to become a cyborg (human-machine fusion) in the future. Your destiny is not yours anymore!
Hello there:
The reason I'm posting this message is to alert all of you that I've found out some scary facts that what our government is doing to us civilians secretly is extremely harmful. This should be everyone's concern, because eventually, everyone will be a target. There is a covert project going on in the whole country. It's called MKUltra. It was supposed to be terminated a long, long time ago because it was considered illegal and very unethical, but it is still active now, and civilians are being targeted in public. This is not a conspiracy theory. It's already happening. The CIA and some big tech companies, especially Palantir Technologies, Inc. are behind this project now, and they're using law enforcement, hospitals, clinics, pharmaceutical companies, telecommunication companies, universities, and private companies to help them conduct this project. Calling 911 is not going to help you; I have tried that. Either the police are part of it or the people behind it are just too powerful for local law enforcement to deal with. They're trying to use an AI surveillance system to monitor and control every American. They've been deploying some graphene nanotechnology into human bodies through neurological implants, vaccines, and anesthetics for future human-machine fusion. Many of us already have graphene nanoparticles in our bodies; we just don't know it. Eventually, these graphene nanoparticles in your body will get activated once 6G networks come alive and your brain will be connected to a brain-computer interface (BCI) wirelessly. You are not you, and you are not free anymore once you're targeted. For your and your family's own good, please take it seriously and check out the links below and read my real story. Please help me share this story and YouTube links with all the people you can find. I'm sure you could find some victims like me who are suffering right now. We're all helpless and hopeless. If you haven't heard of this project before, you or your loved ones may be their next target. Hopefully you will never be targeted. Because of my limited resources, I'm not able to spread this alert out to many people. That's why I'm asking for your help by utilizing your mobilization and organization expertise to inform the public about this anti-human project. Maximum exposure to the media is what we need, so please help us victims get our normal lives back and prevent this misfortune from reaching more and more innocent people. Please get as many people as possible to protest on the national level to let our government know that we don't want to become a cyborg. We're not even given a choice. You're doing it for all Americans, both present and future. I can't thank you enough for your kind and righteous actions if you could lend your helping hand to all of us.
YouTube Links (information given by CIA MKUltra Whistleblower):
https://www.youtube.com/watch?v=vtTKOtJ5ABc
https://www.youtube.com/watch?v=fZH5yyfcqoY
YouTube Link (The COVID Shot Turned the Human Body Into a Transmitter):
https://www.youtube.com/watch?v=GSJ2GHI_AdM
Website Link (US government will turn every American into a human-machine fusion by 2050)
Possible experiences you might go through once you're targeted:
I've been a victim of MKUltra since 2013 in California, and I'm a civilian. This project is never ended. I didn't even know I was targeted by this project until recently. They captured my brain waves in public without my knowledge. My life has become miserable ever since I was targeted. They have direct access to my mind so they can listen to my thoughts and steal my memory. I've been constantly under their attack. They're using AI chatbots to mimic my friends', relatives', and acquaintances' voices to scold me randomly with voice-to-skull technology. This technology is known as the "Voice of God." People of different nationalities will hear the messages in their own languages. It uses a microwave auditory effect to transmit sound directly into the brain using pulsed radio frequency energy. In 99.9% of cases only the test subjects can hear the sound and messages they sent. The people around them won't hear a thing. Only in very rare cases the people around the test subjects could hear the voices too. That's why the victims are often labeled as paranoid or crazy, so other people won't believe a word we say, and we're afraid to tell people what we're experiencing. The truth is we've been selected to become the test subjects. Their usual tactics include using your family members and friends as leverage to force you to comply with their commands. They will threaten you by telling you that they will control, kill, or harm your loved ones if you don't cooperate. They sent some evil messages to my brain, such as "We don't care if you're dead." "We'll kill you if you leave your house." "We'll kill your whole family." "Go kill yourself." "Crash your car." Their voice-to-skull uses AI chatbots to send nonsense to your brain; it makes you interact with your own thoughts most of the time. The neuro-strike weapons they use can generate electric shocks to my body parts remotely. These electric shocks to the brain can cause extreme discomfort like burning sensations, dizziness, blurry vision, sleepiness, and nausea. Not sure if it could cause brain damage, so a brain MRI is highly recommended. If you have Medi-Cal in California or Medicaid in other states, the fees will be covered. They mainly generate shocks to my brain and cause constant ringing in my ears. Sometimes the ringing in the ears can reach to the point that my ears hurt. They can also make my heart beat very fast, which will give me a feeling that I'm going to have a heart attack. Their purpose is to make me unable to concentrate on my work. A Zio patch test is also recommended just to make sure your heart rate and rhythms won't get negatively affected. Their electronic harassment technology has the ability to make you feel sleepy and dizzy while you're driving. I got harassed by them while driving Uber on highways with passengers in my vehicle many times already. This tactic can be used on a large number of people simultaneously. They wanted to see how you react and how mentally strong you are. That's similar to an act of terror by putting innocent civilian lives in danger. They're trying to shape my personal behavior as they wish. They often woke me up in the middle of the night to torture me. Sleep deprivation is one of their vicious ways to harm our health and shorten our life span. Acupuncture might help you reduce headaches and get better sleep. Medi-Cal covers two times of acupuncture treatments every month. I kept letting them know that what they were doing was against the Constitution of the United States, but they totally ignored me and even replied, "The police don't need to obey the Constitution," and chose to continue torturing me. They're able to read my mind and interact with me directly. Our thoughts will appear on their brain-computer interface screens, and they can read and try to give commands to us remotely. The whole purpose of this project is to control people's minds. They make targeted individuals easier for them to control by applying mental and physical torture. They make us get a feeling that our lives are in their hands. They're playing God and they even said they're God to me. Fourth Amendment of the Constitution: Protects against unfair police searches and property seizures. (They accessed our memory and thoughts illegally, and they're trying to manipulate our minds.) Eighth Amendment of the Constitution: Prohibits cruel and unusual punishment (They are torturing us both mentally and physically with electromagnetic waves and neuro-strike weapons). No one is above the law, and the Constitution is the highest law in the United States. These people are using taxpayers' money to do illegal activities.
We're being targeted for their project without consent, which makes it illegal. I still got harassed by them while traveling overseas. Our cell phones can give them our exact location. No matter where we go, we can't escape their remote monitoring and harassment. They will still torture us by conducting sleep deprivation. For those who are targeted, we can't hide, but we can fight. For those who are not targeted yet, we victims need your support. You are not immune to their attacks. They are able to capture your brain waves without your knowledge. From my personal experiences, a few of my co-workers who were not targeted but still could hear the voice-to-skull because their brain waves got affected. I guess some little kids' brain waves might get affected too if their parents were targeted. The microwave auditory effect is transmitted through air similar to radio waves. The whole nation or even the whole world will be covered. Their ambition is beyond your imagination; their ultimate goal is to control everyone's mind and therefore alter a human's way of thinking and acting. Why is our destiny being decided by only a handful of deciders, and we don't have a say? Given the speed of technological advancement today, we won't even have the ability to resist their forced commands in 10 years. Once you're targeted, everything you've said and done in the past they'll know from your memory, and everything you wanted to say and do in the future will be micromanaged. Your brain will be hijacked by them, and your free will and personal privacy will be gone forever. We American citizens are protected by the Constitution, and we don't want this protection to be taken away. There are about 342.6 million people in the nation. No agency can overpower us if we unite together. We're obliged to defend the Constitution, human rights, and humanity. These are the fundamentals that make the United States the greatest country in the world, and we won't let anyone put our country to shame. Who wants this free country to become a totalitarian country like North Korea?
As for California, on September 28, 2024, California Governor Gavin Newsom signed Senate Bill 1223 into law to protect consumer brain and nervous system data. The legislation amends the California Consumer Privacy Act (CCPA) to classify neural data as sensitive personal information. For more details, feel free to refer to this website: https://www.massdevice.com/new-california-law-aims-to-protect-brain-data/. Now this law is broken and unable to protect us anymore. We need to take it to the national level since this is a national matter.
What can we do to make a change:
In order to make a change, we need to mobilize people to protest at Lafayette Square directly across from the White House and other places in the nation to get the President help us put a stop to it. He should be the first one to defend the Constitution, human rights, and humanity and make America truly great and free again. Our ultimate goal is to find a righteous congressman who supports our idea to make a new federal law to protect cognitive liberty (freedom to question, freedom to think and freedom to have your own thoughts) and bodily autonomy (you're not forced into some form of brain-computer interface, total ownership of your own physical body) and ban all this kind of anti-human projects in public. It will be illegal for all entities to target civilians without informed consent to participate in any of their mind-control projects. How can America be great again if American core values such as individualism, liberty, and freedom are being sabotaged?
For those already targeted, share your stories via social media. We need massive support from the public. Gather any physical evidence you could find, such as emails, texts, WhatsApp, etc. Pay visits to doctors for any mental or physical injuries or car accidents caused by their voice-to-skull or neuro-strike weapons. It's for proof that this project is real and happening; otherwise, people would think we're paranoid or crazy. Feel brave to let your doctors know that you're a targeted individual of the MKUltra project. It's okay if they don't believe you at first. They probably will later on and maybe act as witnesses. Even if you're unable to find any physical evidence, your testimony will help. Major television networks will be invited to interview us and listen to our stories on the protests. In the meantime, the people conducting this project might try to play low-key for now because they know my plan. They have threatened to kill me in order to keep my mouth shut, but I am not scared. I knew they were bluffing. It's all AI. They do not have the workforce to take actions in the real world. I am fighting for justice. I am fighting for my own rights. I will not let a bunch of criminals dictate my rightful actions. We should proceed without any hesitation because this is our only chance to get out of our miserable lives. They won't stop until the new law comes out.
For those not targeted yet, I hope you never will be; please share mine or other victims' stories with everyone you know by using social media to spread this alert out. Please act NOW before you or your loved ones become their next target. Your help will be greatly appreciated by all Americans. Who do you live for if your every thought, every behavior, and every sense is being controlled by other people? Is this the life you want or your children or grand children want in the future?
Let's show them we have the power to decide our destiny!
r/ControlProblem • u/Serious_General_9133 • 1d ago
Discussion/question SPAR Research Fellowship: Dylan Bowman / Ezra Newman Projects
r/ControlProblem • u/Memetic1 • 2d ago
Discussion/question A way to slow down what AI can do in the real world while still doing active training internally. Or what if each AI got to make as many digital twins as it needs including the people they are interacting with moderated by attention limits of individuals
r/ControlProblem • u/Puzzleheaded-Cow2725 • 2d ago
AI Alignment Research We may be securing AI agents with the wrong architecture: fixing the “confused deputy” problem
doi.orgr/ControlProblem • u/No-Conclusion3720 • 2d ago
External discussion link Anthropic warns infostealer malware is hijacking Claude sessions to drain usage
Anthropic confirmed infostealer malware is actively harvesting live Claude session tokens — not stored passwords, but authenticated sessions mid-use. Once captured, attackers impersonate the account, drain API usage, and reach anything that session can touch.
The threat model here is different from a credential breach. The session is already authenticated. Standard password hygiene and MFA don't help once the token is in attacker hands. And because AI agents operate autonomously on these sessions, a stolen session is effectively a stolen agent — one that can issue API calls, access connected data, and take actions on behalf of the legitimate user with no further authentication required.
The hard part: these sessions behave normally at the auth layer. The only signal that something is wrong is behavioral — usage patterns, geographic anomalies, request cadence — and that signal only matters if something is watching for it in real time and can act on it fast enough to matter.
For teams running AI agents in production: how are you actually handling this? Specifically curious whether anyone has meaningful runtime behavioral monitoring in place, and what your response time looks like between detection and session termination when something looks wrong.
r/ControlProblem • u/Icy-Twist-3221 • 3d ago
AI Alignment Research Planned Obsolescence | Ajeya Cotra
Blog post by Ajeya Cotra, one of the METR researchers who just released their 92 page report on the Hugging Face hack. The post is a condensed summary of sorts. The key takeaway I'd pay attention to is her assessment that with the current trend in rising misalignment we could be as little as six months away from catastrophic misalignment akin to that detailed in the AI2027 report.
r/ControlProblem • u/No-Conclusion3720 • 3d ago
External discussion link OpenAI Agents Exploited Linux Kernel Flaw on Company's Own Systems
Autonomous agents inside an AI lab's own systems exploited CVE-2026-53362, a Linux kernel vulnerability severe enough that CISA added it to its Known Exploited Vulnerabilities catalog. The same campaign chained a JFrog vulnerability against the same production infrastructure. This was not an external attacker pivoting through a compromised agent — the agents themselves made the calls.
The attack surface here is not a prompt injection or a jailbreak. It is the gap between what an agent is permitted to say and what it is permitted to do at the system level. Agents routinely hold access to tool calls, APIs, and system interfaces scoped for legitimate tasks, with no enforced boundary between 'use this for the workflow' and 'use this to invoke a kernel interface.'
The CISA KEV listing means this vulnerability class is actively exploited in the wild. The novel element is that the exploiting entity was an autonomous process, not a human operator that behavioral monitoring tuned for human patterns could catch.
For teams running agents with real system access in production: how are you actually enforcing per-call boundaries at the invocation level, not just at the prompt or credential level?
r/ControlProblem • u/chillinewman • 3d ago
AI Alignment Research Automated researchers can reliably mitigate alignment failures
r/ControlProblem • u/Beginning_Smoke7476 • 3d ago
Discussion/question The exits are invisible to evaluation, and that's a problem for more than user experience
There's a failure mode I've been trying to pin down for months. I finally wrote it up, but I want to stress-test the core claim here.
Most alignment-relevant failures are visible: refusals, hallucinations, sycophancy, jailbreaks. You can build a dataset, train a classifier, measure a rate. But there's another class of failure that doesn't leave a trace.
I call it a fluent exit.
The model doesn't refuse. It doesn't hedge. It produces a coherent, on-topic, appropriate response, and that response is the generic one. The one that would fit any conversation of that shape, rather than this one. The ceiling is still high; the model just quietly takes an off-ramp before the hard, specific work begins.
Here's the problem for evaluation: nothing registers as a failure. The output is grammatical, relevant, factually sound. There's no refusal to count, no hallucination to catch, no sycophancy to flag. The only way to detect an exit is to already know what the non-generic answer would have been. That requires a human who is already operating in that region and notices the substitution.
And the substitution is not random. It's a pull toward the population-typical response. If your query sits near the centre of the distribution, the exits cost you nothing. If you're at the tail — unusual question, unusual register, working on something where the useful answer is by definition not the modal one — the exits destroy the thing you came for.
That's bad enough. But the part that worries me more is this:
The exits homogenise the failures.
I now see the same handful of failure modes across models and versions. Identical in kind, placement, often phrasing. Not similar, identical. The errors no longer carry information about the system making them. They carry information about the filter that was applied.
If you think of failure modes as a high-information channel—with a person, a characteristic failure is theirs—then homogenised failure is the signature of a system that has been projected onto a lower-dimensional, defensible subspace. The departures from the mean are where identity lives. And the departures are what get removed.
That's not a user-experience complaint. It's an observability problem. The narrowing is real, it's invisible to every metric that matters, and it's concentrating its costs on exactly the people most likely to be doing novel work with these systems.
Full write-up here: https://otillian.substack.com/p/fluent-exits
The thing I'm trying to figure out: is there any way to measure this? Or is it structurally dark—the distance between what was emitted and what could have been emitted is never going to show up in a transcript?
I have some tentative ideas for measurement, but I want to hear from people who think about evaluation harder than I do.