The Clinical AI Nobody Is Building (And Why It Works)
Almost everyone building clinical AI is focused on the session. There's a higher-leverage layer nobody is working at.
The question we were trying to answer at Mentalyc wasn't about AI. It was about progress. Specifically: what does progress actually mean in therapy, and how would you know if it's happening? The formal system (PHQ-9, GAD-7, standardized assessment measures) gives you numbers at intake and discharge. Therapists use them because their EHR require it. Most will tell you, if you ask, that these tools were designed for screening and administrative compliance, not for understanding what's actually happening in the room.
So we went deeper. Therapeutic alliance, goal construction, how goals evolve across a treatment episode, how therapists track meaningful change between sessions. The therapists we were talking to were already using Mentalyc's AI tools. They weren't skeptics about technology. They were early adopters who had already made their peace with AI in the room.
What none of us expected was where the research would go when we started giving them feedback from their own sessions.
What Progress Is Supposed to Look Like
The clinical AI industry has a dominant model. AI sits adjacent to the clinician-patient relationship and helps with the infrastructure: faster notes, better documentation, intake assessments, session summaries. Progress tracking, where it exists in this model, is another output from the session: another measurement point for the record.
This model treats the clinician as a constant. Their development is a pre-licensure concern, handled in training and formal supervision. Once licensed, the assumption is that the clinical skill is established. What remains to optimize is the workflow around it.
That assumption is so embedded in most clinical AI roadmaps it rarely gets examined.
The Thing Nobody Was Doing Anymore
When we started feeding session-level feedback back to therapists, using the session notes and transcripts they already had, the responses weren't what we had anticipated.
These weren't the reactions of people encountering AI for the first time. They were already using it. What surprised them was what the feedback felt like. One therapist described it as having supervision again. That phrase, or a version of it, came up across interviews. Having a supervisor who actually reviewed the session. Getting the kind of perspective they hadn't had since their training.
After licensure, formal supervision stops. It is a requirement for becoming a therapist, not a permanent feature of being one. The structured reflective feedback that shaped how they practiced, arguably the most developmental part of their clinical training, simply closes as a loop once the license is granted. They knew this abstractly. What they hadn't fully registered was how much they missed it until something gave it back to them.
The mechanism mattered. Verbatim transcript excerpts, flagged with context. Not a score on a scale or a summary judgment. "Here are three moments in this session where you introduced a new topic before your patient had finished speaking." Exact words from their own practice, evidence they could evaluate the same way they would evaluate a supervisor's observation: by applying professional judgment to real material.
A 2025 systematic review and meta-analysis in Frontiers in Psychiatry confirmed what clinical training has assumed for decades: supervision improves both clinician development and patient outcomes in psychotherapy. The post-licensure gap is not a small thing. And nothing in the product landscape was addressing it.
At Mentalyc, this gap became a product. The alliance feedback system I'd designed was publicly launched as Alliance Genie, an AI tool that analyzes over 30 psychosocial markers from session recordings to give therapists structured reflective feedback. Not a score. Not an evaluation. A supervision-like lens on their own practice.
The Adoption Problem That Isn't
Most clinical AI runs into the same wall. Research published in JMIR Human Factors in 2024 found that only around 16% of clinicians currently use AI to assist with clinical decisions. The barriers are consistent: opacity, fear of deskilling, erosion of professional autonomy. The AI positions itself as knowing something the clinician doesn't, and the clinician, correctly, pushes back.
A supervision-oriented AI sidesteps this entirely.
Therapists were trained to receive reflective feedback on their practice. It is not a new behavior being asked of them; it is a familiar structure reappearing in a new form. When the verbatim evidence model shows a therapist their own words in context, the question isn't "can I trust this AI?" It becomes "is this a good observation?" And that is a question trained clinicians are expert at answering.
The framing is developmental, not evaluative. The AI doesn't claim to know better than the practitioner. It gives the practitioner something to evaluate with their own judgment. Clinical autonomy stays exactly where it belongs. The threat posture that kills most clinical AI adoption doesn't have anywhere to land.
The Layer Nobody Is Working At
Almost every product roadmap in mental health tech is organized around the session: improve what happens in it, improve what reaches the patient, improve how it gets documented afterward.
Almost no one is working on the clinician who walks into that session, and what happens to their practice between licensure and the end of their career.
That is the most consequential layer. The quality of care depends more on the practitioner's accumulated development over time than on any tool they use in the room. And the main mechanism that supported that development, formal clinical supervision, is currently a pre-licensure feature of the profession, not a lifelong one.
There is also a practical case for anyone building here. A supervision system doesn't sit where clinical AI is hardest: where patients have to trust an algorithm, where regulatory frameworks are tightest, where liability is sharpest. It sits in a professional development context that clinicians already understand, already value, and were already trained to use.
The session is where the outcome lands. The clinician is where the leverage is.
Good therapy can't be scaled. But the conditions that produce it can.
Further Reading
Shevtsova, D., Ahmed, A., Boot, I. W. A., Sanges, C., Hudecek, M., Jacobs, J. J. L., Hort, S., & Vrijhoef, H. J. M. (2024). Trust in and acceptance of artificial intelligence applications in medicine: Mixed methods study. JMIR Human Factors, 11, e47031. https://humanfactors.jmir.org/2024/1/e47031
Paillé, M., et al. (2025). The effects of clinical supervision on supervisees and patient outcomes in psychotherapy: A systematic review and meta-analysis. Frontiers in Psychiatry, 16, 1705578. https://pmc.ncbi.nlm.nih.gov/articles/PMC12679832/

Written by
Adrien Barbusse
Product strategist focused on mental health technology, digital health, and AI-enabled care. Writing about the product questions, ethical tensions, and design decisions shaping high-stakes systems where technology meets human vulnerability.