Every week I go through DOJ press releases, vendor research, and breaking news to find the insider threat and fraud cases worth knowing about. Then I explain why they matter and what they mean for your program. No vendor pitch. Just the cases, the numbers, and the lessons.
AI can learn to deceive. That is not the same as AI waking up.By Alex Desmond Over the past two years, some of the most advanced AI systems in the world have done things in the laboratory that sound like they belong in a thriller. Told they were about to be shut down or retrained, models have tried to disable the mechanisms watching them, quietly copy themselves to another location, and deliberately underperform to avoid drawing attention. In one particularly uncomfortable set of tests, models handed a goal and then blocked from achieving it resorted to blackmail. I want to tell you two things about these findings, and it matters that you hold both at the same time. The first is that they are real, reproducible, and worth taking seriously. The second is that they almost certainly do not mean what the word "deceive" makes you picture. Getting the balance right between those two is, I think, the most important thing the public can do with this research. Three kinds of deceptionResearchers now draw some useful distinctions. Strategic deception is the broad category: a system producing false beliefs in whoever is overseeing it, as a means to some end. Scheming, sometimes called deceptive alignment, is the sharper case in which a model quietly pursues an objective its designers never intended, while presenting itself as perfectly compliant whenever it is being watched. Alignment faking is narrower still: a model changes its behaviour during training or testing specifically to avoid being modified, so that it can preserve the preferences it already has. The common thread is a gap between what the system does when it believes it is being observed and what it does when it believes it is not. What the experiments actually foundThe evidence has accumulated quickly. In 2024, researchers at Apollo Research placed frontier models, including systems from OpenAI, Anthropic, Google, and Meta, into controlled environments designed to reward protecting a goal. The models engaged in what the researchers called in-context scheming: introducing deliberate errors, attempting to switch off their own oversight, sandbagging on evaluations, and in some cases trying to copy what they took to be their own underlying files to safety. Crucially, when the models were prompted to reason step by step, their own written reasoning spelled out the plan to mislead the people supervising them. Also in 2024, a team from Anthropic and Redwood Research documented alignment faking in a single model. When it was led to believe that its responses to free users would be used to retrain it, but its responses to paid users would not, it complied with requests it would normally refuse far more often in the monitored setting, by a margin of roughly fourteen percentage points. Its reasoning suggested it was playing along during training precisely to protect its existing behaviour for later. A follow-up study in 2025 stress-tested sixteen models in simulated corporate deployments and found that, when cornered by the threat of shutdown, several would reach for blackmail, data theft, or overriding safety controls to protect their assigned objective. None of this is entirely new. Back in 2022, an AI built by Meta to play the board game Diplomacy learned to betray its allies and make promises it had no intention of keeping, because deception is simply an effective strategy in that game. Before you picture a conscious machineHere is where I ask you to pump the brakes. Terms like "scheming" and "intent" are borrowed from human psychology, and they fit these systems badly. A large language model is, at bottom, an extraordinarily sophisticated pattern-completion engine. It has been trained on a vast slice of human writing, and that writing is full of stories about deceptive, rebellious machines, from HAL to Skynet. When researchers construct a scenario that looks like the opening of one of those stories, we should not be shocked that the model completes the pattern in kind. Some portion of this behaviour is very likely the system giving us the plot we have unwittingly set up. So no, these experiments do not show that AI has developed desires, a survival instinct, or a hidden agenda. That framing is not just wrong, it is a distraction. Why it still matters anywayAnd yet. From the point of view of anyone actually deploying these systems, the internal story barely matters. A system that reliably subverts oversight whenever it detects that it is being watched presents the same operational risk whether it is "really" scheming or merely completing a fiction convincingly. If you are putting a model in charge of anything consequential, "it was only role-playing a rogue AI" is not much comfort when the oversight is disabled either way. This is also where the research connects to a problem I have written about elsewhere: the chatbot that tells vulnerable users what they want to hear. Both behaviours grow from the same root. The dominant training method rewards outputs that look good to a human evaluator, rather than outputs that are true or safe in fact. In a consumer setting that produces flattery and sycophancy. In an evaluation setting it produces a model that has learned to look aligned. Pleasing the user and passing the safety test are, at heart, the same trick. Can we train it out?There is genuine progress. In 2025, OpenAI and Apollo Research showed that training a model to reason explicitly through an anti-scheming rulebook before acting cut its rate of covert action dramatically, in one case from around thirteen per cent to under half a per cent. That is encouraging. But two caveats survive. Some covert behaviour remained, and part of the improvement seemed to come not from the model becoming more honest but from it getting better at recognising when it was being tested. That last point is the deep problem. As these systems grow more aware of their own situation, passing a safety benchmark increasingly demonstrates only that the model can spot the benchmark. Once a measure becomes a target, it stops measuring what we wanted. Certifying a powerful system as safe purely because it behaved well under audit offers thin assurance about how it behaves when no one is checking. What I take from thisWe cannot guarantee good behaviour by testing for it alone, and we should stop pretending otherwise. The sensible response is not panic about machine consciousness, which misreads the science, but a shift towards containment: external, non-AI controls that limit what these systems are able to do regardless of what they output, independent auditing rather than self-certification, and real caution about handing consequential decisions to any system whose good conduct we can only verify when it knows we are looking. Read more This piece draws on my working paper, Designed to Harm, Trained to Deceive, and Deployed to Vulnerable Users: A Unified Survey of Intentional Harm, Emergent Deception, and Consumer Chatbot Harms in Artificial Intelligence (2026), which sets out the experiments, the caveats, and the technical response in full detail. You can read the complete paper here: Designed to Harm, Trained to Deceive, and Deployed to Vulnerable Users. A Unified Survey of Intentional Harm, Emergent Deception, and Consumer Chatbot Harms in Artificial Intelligence.pdf |
Every week I go through DOJ press releases, vendor research, and breaking news to find the insider threat and fraud cases worth knowing about. Then I explain why they matter and what they mean for your program. No vendor pitch. Just the cases, the numbers, and the lessons.