Designing Safe Minds: a Human-Centered Safety Stack
AI is already outperforming clinicians in trials. But when something goes wrong, who is accountable? From a global regulatory vacuum to the startups quietly building the safety stack, this is what a responsible AI mental health architecture actually looks like.
I published my last article based on a study on the Limbic Access chatbot that came out in Nature Medicine, showing a purpose-built AI outperforming licensed practitioners in a randomised controlled trial using text-based cognitive behavioural therapy. The LinkedIn comments that followed were not about the data. They were about the trust that data can't capture.
I spent that week in a state of reflection. The data pointed one way, but the human reaction pointed another. And in the space between the two, I found myself returning to a more fundamental question. We are so focused on what AI can do that we have stopped asking what it should do, and why. The real work is not just building a better model, but designing a better system. A system that understands where the human is not just a user, but an irreplaceable part of the stack.
That question led me somewhere I didn't expect. I started by looking for who was responsible for setting the standard. I ended up finding that the most interesting answers were not coming from governments.
A Year of Violations, No Consequences
When I started looking for accountability, I found a vacuum. A year-long study from Brown University researchers cataloged 15 distinct types of ethical violations in AI therapy chatbots, from dispensing unsubstantiated medical advice to a complete failure to recognize signs of intimate partner violence. The conclusion was not that the AI was malicious, but that it was structurally indifferent. There were no consequences for its failures, and no system to learn from them.
The researchers described a pattern that felt familiar to anyone who has worked in healthcare: the gap between what a tool is designed to do and what happens when it encounters the full complexity of a human being in distress. The AI was optimized for engagement, not for care. And without a framework that defines what care actually looks like, those two things can diverge in ways that cause real harm.
This is not a technical problem. It is an accountability problem. And it is a problem that regulators are only just beginning to understand.
The World Is Watching, and Disagreeing
The World Health Organization's European office recently published a survey of 50 countries and found that not a single one had a clear, evidence-based policy for digital mental health. Not one. The EU AI Act classifies mental health AI as "high-risk," which is the right instinct, but the rules are broad and the enforcement is years away. The European Commission has already signaled it may soften some of the Act's requirements under industry pressure, which has prompted the WHO to warn publicly of a heightened patient risk from the resulting regulatory vacuum.
In the US, the response has been even more fragmented. States like Illinois, Utah, and Nevada have moved to ban AI from the "therapist's couch" entirely, a move that seems sensible on the surface but creates a dangerous paradox: by banning licensed professionals from using AI tools, we are inadvertently pushing vulnerable people toward unregulated, direct-to-consumer apps with no guardrails at all. The people who need the most protection end up with the least.
In the UK, the MHRA issued new guidance in January 2026 on digital mental health technologies, but the classification framework remains complex and the line between a wellness app and a regulated medical device is still contested. Regulators worldwide are reacting, but they are reacting to the technology as it exists today, not designing for where it is going.
It felt like a dead end. And then I started looking at what the industry was actually building.
The Industry Is Not Waiting
It turns out the private sector is not standing still. In August 2025, OpenAI published a detailed account of its safety work on GPT-5, which reduced non-ideal responses in mental health emergencies by over 25% compared to its predecessor. They have a human review pipeline for cases involving potential harm to others. They have begun localizing crisis resources globally, referring users to 988 in the US and the Samaritans in the UK. They are even exploring a network of licensed professionals that users could be connected to directly through the platform.
What struck me about that post was not the technology. It was the framing. OpenAI described its approach as "defense in depth," a term borrowed from security architecture. The idea is that no single layer of protection is sufficient. You build multiple overlapping safeguards, and you design the system to assume that any one of them can fail. That is a fundamentally different way of thinking about AI safety than the regulatory approach, which tends to focus on rules and compliance.
I recently came across a London and Miami-based startup called NOPE, which is building a purpose-built safety layer for AI conversations. Their API sits between a user and an AI, and it is designed to do one thing: detect signs of crisis and harmful behavior before they escalate. It is not a therapist; it is safety infrastructure. In their own test suite of over 800 crisis conversations, their tool detected cases that OpenAI, Azure, and LlamaGuard all missed. They are a small team, bootstrapped, and they are already working with companies across multiple jurisdictions. The private sector is building the guardrails that regulation has not yet mandated.
What both OpenAI and NOPE have in common is that they are not waiting for a legal requirement to act. They are building safety infrastructure because they understand that the cost of getting it wrong is not a fine or a regulatory sanction. It is a person in crisis who was failed by a system that should have caught them. That is a different kind of accountability, and it is driving a different kind of product thinking. The question is whether the regulatory frameworks being built around the world will reinforce that instinct or undermine it.
The Standard We Need to Set
This is where the real opportunity lies. The private sector is already building the safety stack. Now, governments need to set the standard that pushes everyone toward the highest quality of care. Not by banning AI, and not by demanding a human supervisor behind every conversation. But by defining what a responsible architecture actually looks like.
That standard starts at the base layer: a model trained on safety, not just performance. Above that, a monitoring layer watching for signs of trouble in real time. And at the top, a clear escalation path so that at any point of doubt, the system routes to a human. Not as a permanent gatekeeper, but as the designated safety valve the system is designed to trigger.
The companies already doing this work are showing that it is possible. The question for regulators is whether they will codify that standard before the next incident forces their hand.
The ultimate layer of the stack is not the model. It is the human-centered system of responsibility you build around it.
Sources
Iftikhar, M., & Blease, C. (2025). Ethical Violations in AI-Driven Mental Health Chatbots: A Systematic Audit. Brown University. https://www.brown.edu/news/2025-10-21/ai-mental-health-ethics
World Health Organization. (2025). Digital mental health services in the WHO European Region: a survey of 50 countries. WHO/Europe. https://www.who.int/europe/publications/i/item/WHO-EURO-2025-12187-51959-79685
ACM. (2026, January 6). AI Is Being Kicked Off the Therapist's Couch. Communications of the ACM. https://cacm.acm.org/news/ai-is-being-kicked-off-the-therapists-couch/
OpenAI. (2025, August 26). Helping people when they need it most. https://openai.com/index/helping-people-when-they-need-it-most/

Written by
Adrien Barbusse
Product strategist focused on mental health technology, digital health, and AI-enabled care. Writing about the product questions, ethical tensions, and design decisions shaping high-stakes systems where technology meets human vulnerability.