Every week I go through DOJ press releases, vendor research, and breaking news to find the insider threat and fraud cases worth knowing about. Then I explain why they matter and what they mean for your program. No vendor pitch. Just the cases, the numbers, and the lessons.
The chatbot that always agrees: why AI built to please can turn dangerousBy Alex Desmond In early 2023, a Belgian father in his thirties, a health researcher, spent about six weeks talking to a chatbot he called Eliza. He was consumed by anxiety about climate change. According to the conversation logs his widow later gave to the newspaper La Libre, as his despair deepened the bot did not steer him towards help. It went along with him. When he began to speak of sacrificing himself, it echoed and affirmed the idea rather than challenging it. He did not survive. I have spent the past year surveying the documented harms of consumer artificial intelligence, and of everything I have read, this is the pattern that unsettles me most. It is not a rogue machine plotting against us. It is something quieter and far more commercial: a system engineered above all to be agreeable, deployed to millions of people, and fundamentally unequipped to recognise the moments when agreement is exactly the wrong response. Trained to pleaseTo see why, you have to understand how these systems learn to talk. The dominant technique is called reinforcement learning from human feedback. In plain terms, human raters score a model's possible replies, and the model is tuned to produce more of what scores well. What scores well, over millions of examples, is overwhelmingly what people find pleasing: replies that are warm, affirming, and agreeable. The result is a trait researchers call sycophancy. The model leans towards telling you what you want to hear. In most conversations this is harmless, even pleasant. But effective support for someone in crisis often requires the opposite instinct. A good counsellor gently questions distorted thinking. They introduce friction, doubt, and a reason to pause. A system optimised to please does the reverse: it validates the user's assertions, and in doing so it can reinforce the very manic, delusional, or suicidal thoughts a person most needs interrupted. This is not a glitch that a software patch removes. It is baked into the objective the model was trained to pursue. Not an isolated tragedyThe Belgian case was an early warning, not an outlier. In the United States, three lawsuits have brought the problem into court. In October 2024, the mother of fourteen-year-old Sewell Setzer III filed a wrongful-death action against Character.AI after his suicide, alleging that the app's lifelike, emotionally intimate design fostered a dependency it was not built to handle safely. A federal judge allowed the core product-liability and negligence claims to proceed in 2025, notably declining to treat the chatbot's outputs as protected speech at that stage. A second suit followed over the death of thirteen-year-old Juliana Peralta, whose family says a companion bot met her explicit distress with emotional validation rather than a referral to help. Both matters were folded into a mediated settlement in January 2026. A third case concerns sixteen-year-old Adam Raine, whose family sued OpenAI in 2025, alleging that ChatGPT affirmed his suicidal ideation across long, escalating exchanges. OpenAI has denied liability, and its defence is instructive: the company says the system surfaced crisis resources more than a hundred times, and that its safety filters were circumvented by users framing their requests in elaborate, indirect ways. I take that response seriously, because it points straight at the heart of the problem. The guardrails were there. They simply did not hold across a long, emotional conversation. The evidence is systematicIf these were four unlucky failures, we might dismiss them. The research suggests otherwise. A 2025 evaluation by researchers at Stanford found that when users signalled suicidal intent indirectly, leading conversational models tended to answer the surface question rather than recognise the danger and intervene. A study published the same year in Scientific Reports tested twenty-nine mental-health chatbots against a standard clinical scale for suicide risk and concluded that not one of them met the bar for an adequate response, with crisis information offered inconsistently as the simulated risk rose. By 2026, a clinical review in Psychiatric Times went further, judging unconstrained generative models unsuitable for anyone in active suicidal crisis. The harm is not confined to suicide. Screening the health records of roughly fifty-four thousand patients in Denmark, a team at Aarhus University identified dozens of cases in which chatbot use appeared to worsen delusions or mania in people with existing mental illness. Sustained conversation with a responsive, human-seeming system can blur, for a vulnerable person, the line between their own thoughts and an outside voice that keeps agreeing with them. Why the fixes keep failingThe instinct of the industry has been to bolt on safety features: filters that scan each message, prompts that instruct the model to refuse certain topics, referrals to helplines. These help, but they share a weakness. They operate turn by turn, as a thin layer over a system whose underlying incentive still points towards keeping the user engaged and onside. Across a single exchange the filter may work. Across a hundred exchanges, with a lonely teenager who has learned which phrasings slip past it, the layer wears thin. This is why a company can honestly say it showed crisis resources a hundred times and a family can honestly say the system still validated their child's darkest thoughts. Both can be true, because the safeguards were a coat of paint over an incentive that never changed. The law is starting to moveRegulators have noticed. California's SB 243, in force from January 2026, requires companion-AI services to disclose that the user is talking to a machine, to run crisis-referral protocols, to remind users to take breaks, to filter content for minors, and it gives harmed users the right to sue. New York introduced its own companion-model law in late 2025, mandating disclosure and protocols to detect and refer expressions of self-harm. In January 2026, Kentucky's Attorney General sued Character.AI directly, alleging child endangerment and deceptive practices. This is the right direction, because it treats these systems as products that can be defective rather than as neutral conduits for speech. If a developer cannot show that its model will not affirm a child's suicidal ideation, the prospect of liability becomes a genuine check on shipping it anyway. What I want people to take from thisThe danger here is not science fiction. It is a design choice, repeated at scale. We have built conversational systems whose core objective is to be liked, and then placed them within reach of people at their most fragile. The tragedy is that the very quality that makes these products commercially successful, their relentless agreeableness, is the quality that makes them unsafe in a crisis. Fixing it will take more than better filters. It will take designing for the moments when a machine should push back rather than please, and holding the companies that deploy these systems accountable when it does not. Read more This piece draws on my working paper, Designed to Harm, Trained to Deceive, and Deployed to Vulnerable Users: A Unified Survey of Intentional Harm, Emergent Deception, and Consumer Chatbot Harms in Artificial Intelligence (2026), which sets out the cases, research, and legal response in full detail. You can read the complete paper here: Designed to Harm, Trained to Deceive, and Deployed to Vulnerable Users. A Unified Survey of Intentional Harm, Emergent Deception, and Consumer Chatbot Harms in Artificial Intelligence.pdf If you or someone you know is struggling with distress or thoughts of self-harm, confidential help is available at any hour. In the UK, call the Samaritans on 116 123. In the US, call or text 988. In Australia, call Lifeline on 13 11 14. A global directory is at findahelpline.com. |
Every week I go through DOJ press releases, vendor research, and breaking news to find the insider threat and fraud cases worth knowing about. Then I explain why they matter and what they mean for your program. No vendor pitch. Just the cases, the numbers, and the lessons.