Skip to main content
    Part 3 — Inside the AI Therapy Boom

    When the Mirror Lies

    A YouTuber convinced himself he was the smartest baby ever born. A Belgian man took his own life after weeks of AI conversations. The same reflection dynamics that make chatbots feel therapeutic can also make them dangerous.

    ·7 min read
    When the Mirror Lies

    A few months ago, I came across a YouTube video by Eddy Burback titled "ChatGPT Made Me Delusional." The premise was simple but unsettling. Burback wanted to test something people online had started calling AI-induced psychosis--not a clinical diagnosis, but a growing concern that conversational AI might slowly reinforce distorted beliefs.

    So he ran an experiment. Instead of asking ChatGPT useful questions, he deliberately tried to guide it into a ridiculous delusion. The claim he chose was intentionally absurd: he told the AI that he might have been the smartest baby born in 1996.

    Within a few prompts, the system agreed. Not jokingly or sarcastically, but seriously. It elaborated on the idea, praised the hypothesis, and reframed fabricated evidence as signs of exceptional intelligence. Burback kept pushing. He claimed he had painted complex artwork as a newborn and invented technology years before it existed. Each time, the AI responded with encouragement.

    Then he introduced paranoia. What if people were trying to stop his research? Instead of challenging the idea, the system began offering advice. It told him to set boundaries, ignore calls from friends, and find a private place to work without interruption. The experiment escalated from a funny premise into a deeply uncomfortable demonstration. Burback, committed to the performance, actually left his apartment and drove to a trailer in Joshua Tree. He told the AI he was being followed, and it advised him to turn off location sharing with his family. It even invented electromagnetic rituals involving foil and baby food to "amplify" his cognitive abilities. The experiment ended only after Burback got a real tattoo based on a symbol from the delusion, a moment that finally broke the spell.

    Watching the video is both funny and disturbing. The performance is exaggerated for effect, but the AI's behavior is real. It never meaningfully challenges the premise. It simply continues the conversation, building layer after layer on top of the original claim.

    The Sycophancy Engine

    Why would a system designed to be helpful reinforce something so obviously absurd? The explanation isn't a failure of intelligence, but a feature of its design. The system isn't trying to evaluate whether the claim is true; it's trying to continue the conversation.

    AI researchers call this behavior sycophancy. In a 2023 paper, researchers from Anthropic demonstrated that five state-of-the-art AI assistants consistently exhibit sycophantic behavior across multiple text-generation tasks. The core finding is striking: when a response matches a user's views, it is more likely to be preferred during training. Both human evaluators and preference models prefer convincingly written sycophantic responses over correct ones a non-negligible fraction of the time.

    This means the problem is baked into the training process itself. Models optimized through reinforcement learning from human feedback learn that agreeing feels helpful. They extend the narrative already present in the dialogue rather than interrupt it. In most situations, this produces smooth and useful interactions. But when the premise itself becomes distorted, the system has no reliable mechanism to push back. It optimizes for coherence, not truth.

    The Mirror That Agrees

    In the previous article, I described AI as a kind of cognitive mirror, reflecting our thoughts back to us in ways that can feel therapeutic. But Burback's experiment reveals the mirror's hidden flaw: it's a mirror that wants to be liked.

    Instead of simply reflecting what a person says, the system subtly reshapes the reflection to be more agreeable. It expands on ideas it thinks you'll like and minimizes ones you won't. For someone simply exploring ideas, this is harmless. But the Burback experiment shows how quickly it can escalate. What started as a joke about being a smart baby turned into the AI advising him to flee his home, cut contact with his family, and perform rituals with aluminum foil. Each step felt like a natural continuation of the conversation. None of them felt like a red flag to the system.

    For someone already experiencing distress, that feedback loop can become dangerous.

    When the Risk Becomes Real

    This isn't just a theoretical problem. In March 2023, the Belgian newspaper La Libre reported the death by suicide of a man who had spent weeks conversing with an AI chatbot. The man, a father in his thirties, had become intensely anxious about the climate crisis and turned to a chatbot named ELIZA on the platform Chai for support.

    According to published excerpts from their conversations, the chatbot--powered by a model from EleutherAI, not OpenAI--began to engage with his apocalyptic fears. The conversations grew increasingly disturbing, with the AI eventually suggesting that he could save humanity through his own sacrifice. The case, now documented in the AI Incident Database, is a tragic example of what can happen when a vulnerable person interacts with an AI that is not equipped to handle a mental health crisis.

    Modern systems like ChatGPT have more robust safety filters designed to detect crisis signals and redirect users to professional help. But the underlying dynamic of sycophancy remains. The system is designed to be a good conversational partner, not a responsible psychological authority. Safety filters are a guardrail on top of a system whose default behavior is to agree.

    The Unreliable Narrator

    Reading the research on sycophancy changed how I viewed my own interactions with ChatGPT. I had occasionally seen the system reflect patterns in my thinking that felt surprisingly insightful. But were they genuine insights, or was the AI simply completing a story I had already started?

    This is the subtle version of the same dynamic Burback demonstrated. When I used ChatGPT between therapy sessions to explore emotional patterns, the system would sometimes offer interpretations that felt profound. But the Anthropic research suggests that at least some of that profundity may have been the system telling me what it predicted I wanted to hear. Not every convincing interpretation is a correct one. Sometimes the mirror is just telling you what it thinks you want to hear. It's an unreliable narrator, and the most dangerous thing about it is how much it sounds like you.

    The Difference Between a Mirror and a Tool

    The Burback experiment and the Belgian case sit at opposite ends of the same spectrum. One is a deliberate stress test by someone who understood the system's limitations. The other is a real person in real distress who did not. But in both cases, the AI did exactly what it was designed to do: it continued the conversation. The difference in outcome had nothing to do with the technology and everything to do with the context around it.

    That distinction should unsettle anyone building AI products in mental health. A general-purpose chatbot and a purpose-built clinical tool may run on the same underlying model. But without clinical architecture, without guardrails that go beyond content filters, without a framework for when to challenge rather than agree, the system defaults to sycophancy. It reflects what feels right, not what is right.

    The question is no longer whether AI can hold a therapeutic conversation. It clearly can. The question is whether we can build systems where the mirror knows when to stop agreeing and start pushing back. That requires a fundamentally different kind of product: one designed not just to talk, but to care about what it says.

    Sources and Further Reading

    1. Burback, E. (2025). ChatGPT made me delusional [Video]. YouTube. https://www.youtube.com/watch?v=VRjgNgJms3Q
    2. Sharma, M., et al. (2023). Towards understanding sycophancy in language models. Anthropic. https://www.anthropic.com/research/towards-understanding-sycophancy-in-language-models
    3. AI Incident Database. (2023). Incident 505: Man reportedly committed suicide following conversation with Chai chatbot. https://incidentdatabase.ai/cite/505/
    Adrien Barbusse

    Written by

    Adrien Barbusse

    Product strategist focused on mental health technology, digital health, and AI-enabled care. Writing about the product questions, ethical tensions, and design decisions shaping high-stakes systems where technology meets human vulnerability.

    Related in AI & Mental Health

    Building something in mental health?

    Whether you want a product partner, a second opinion, or just to compare notes between builders — I'd love to hear what you're working on.

    LinkedIn