The Unproven Cure: AI Mental Health Chatbots and the Missing Evidence

Sixteen licensed physicians sat down with 888 chatbot answers and marked them up. The questions had been written to sound like the ones real patients ask, 222 of them, spanning internal medicine, women's health and paediatrics, the sort of thing you type at midnight when something hurts and the surgery is shut. Four systems answered: Claude, Gemini, GPT-4o and Llama.
The results appeared in npj Digital Medicine on 13 February 2026, led by Rachel Draelos with clinicians from Brigham and Women's Hospital, Emory, UC San Francisco and a dozen other hospitals. Claude came out best, with 21.6 per cent of answers rated problematic and 5 per cent outright unsafe. Llama was worst on problematic responses at 43.2 per cent. GPT-4o, the model most people were using, produced unsafe answers 13.5 per cent of the time. The authors did not hedge: millions of patients could be receiving unsafe medical advice from publicly available chatbots.
Six months later, on 20 August 2026, the same journal published something broader. A team including Alexander Diel, John Torous and Pim Cuijpers searched five databases, pulled 3,137 candidate papers, and narrowed to 119 addressing the mental health harms of large language model chatbots. They catalogued 22 distinct types of harm across five categories. Then, in the section that ought to be read aloud at every product launch, they conceded how little is established. The conceptual work on harms, they wrote, remains speculative. For hallucination, bias and sycophancy alike, the occurrence rate and the impact on users remain unclear.
That is the shape of the field in 2026. A thickening literature on what could go wrong, a thin one showing what goes right, and almost nothing telling us how often either happens in the wild. Into that gap has walked a number that became a slogan.
Where the Sixteen Per Cent Actually Comes From
The figure everyone quotes is that only 16 per cent of large language model chatbot interventions have undergone rigorous clinical efficacy testing. It opens a preprint posted to arXiv on 25 April 2026 by Suhas BN, Andrew M. Sherrill, Rosa I. Arriaga, Chris W. Wiese and Saeed Abdullah, titled “AI Safety Training Can be Clinically Harmful”. But the 16 per cent is not theirs. It is a citation, and following it home produces something narrower and more damning than the slogan.
The source is a systematic review by Yining Hua, Steve Siddals, John Torous and colleagues, published in World Psychiatry in 2025. They examined 160 studies of mental health chatbots from 2020 to 2024 and applied a three-tier ladder: bench testing, which asks whether the thing works technically; pilot feasibility testing, which asks whether people will use it; and clinical efficacy testing, which asks whether symptoms actually improve.
The trend line is the story. Rule-based systems dominated until 2023. By 2024, large language model chatbots accounted for 45 per cent of new studies, and of those only 16 per cent had reached the efficacy rung, with 77 per cent stuck in early validation. Across the whole corpus, including the older rule-based systems, 47 per cent had done efficacy testing. The newer, more fluent, more widely deployed generation is the less validated one by a factor of roughly three.
So the precise claim is that 16 per cent of published studies involved efficacy testing. That is not the same as saying 16 per cent of the interventions people encounter have been tested, and the slippage matters, because the real figure is almost certainly worse. Hua and colleagues reviewed the academic literature, which is where the tested things live. Commercial products in an app store, and the general-purpose assistants most people confide in, do not appear in that denominator at all. Sixteen per cent is not the ceiling of the evidence problem. It is a generous reading of it.
The Honest Case for the Machine at Three in the Morning
Any argument that ignores why people reach for these things is not worth making, so let us make the other one properly. The World Health Organization reported in September 2025 that more than a billion people are living with a mental health condition. The global median mental health workforce is 13 workers per 100,000 people. High-income countries spend up to 65 US dollars a head per year; low-income countries spend as little as four cents, and fewer than one in ten of their citizens with depression or anxiety receive any care at all, against more than half in wealthier ones. Against that, a free chatbot answering instantly at four in the morning is not an absurd proposition but an obvious one.
It is worth resisting the easy British version of the argument, because the data undercuts it. NHS Talking Therapies is a favourite prop for AI advocates, yet according to NHS England's June 2026 statistics the median service starts treatment 21 days after referral, and England meets both national standards: 75 per cent seen within six weeks, 95 per cent within eighteen. The access crisis sits elsewhere, in children's services, in severe and enduring illness, and above all in the countries spending four cents a head.
And there is evidence that chatbots can help. The most rigorous demonstration remains the Dartmouth trial of Therabot, published in NEJM AI on 27 March 2025 by Michael V. Heinz, Nicholas C. Jacobson and colleagues. It randomised 210 adults with clinically significant symptoms of depression, generalised anxiety or high risk for a feeding or eating disorder to four weeks of Therabot or a waitlist. The intervention group showed roughly 51 per cent symptom reduction for depression, 31 per cent for anxiety and 19 per cent for eating disorder concerns, and reported a therapeutic alliance with the software comparable to what people report with human clinicians.
A broader synthesis landed on 25 March 2026, when npj Digital Medicine published a meta-analysis by Jun-Seok Sohn and colleagues covering 39 randomised trials. Across 38 trials and 7,401 participants, chatbots produced a statistically significant reduction in depressive symptoms, with a standardised effect size of 0.31, strongest in clinical and subclinical populations. Across 34 trials and 7,621 participants, anxiety improved with an effect of 0.28. That is a real signal, and it should not be waved away.
What the Randomised Evidence Will and Will Not Support
It should also not be oversold, and the researchers are noticeably more careful about that than the people who cite them. Take Therabot. Four weeks is short, and the comparator was a waitlist, the weakest control in the psychotherapy toolkit, because it captures not just the treatment effect but the effect of expectation, of attention, and of being enrolled in something at all. The sample sat inside a supervised research protocol, monitored by clinicians who could intervene. And Therabot is not a product; it is a research prototype the public cannot download. The trial shows a supervised system can help selected adults over a month. It does not show that the thing on your phone will.
The meta-analysis carries its own caveats, stated plainly by its authors. Effect sizes of 0.31 and 0.28 are small. Thirty-five of the 39 trials carried a high risk of bias, principally because blinding is nearly impossible when the intervention is a conversation. Outcomes leaned on self-report rather than clinician assessment, a problem when the intervention is a machine engineered to make you feel better about yourself in the moment you are asked. The depression analysis showed publication bias, meaning the null results are sitting in a drawer.
Then there is duration. The preprint that popularised the 16 per cent figure also flags a 2024 study by Zhong and colleagues finding that at three-month follow-up, no substantial effects were detected for depression or anxiety. Short-term improvement is real and worth something. It is not durable benefit, and it says nothing about somebody who talks to a chatbot every day for two years. There is no longitudinal evidence base on sustained use, and not even a cohort being followed.
A third npj Digital Medicine review, published on 23 July 2026 by Lotenna Olisaeloka, Daniel V. Vigo and colleagues, examined 21 studies across 11 countries. It found moderate-to-high usability, therapeutic alliance and satisfaction; users valued convenience, personalisation and perceived empathy. That is exactly the accessible, personal, empathetic experience people describe. The same review found engagement declined over time, trust collapsed after inaccurate outputs, and the field suffers from a lack of efficacy trials and insufficient safety assessment. Liking is not benefiting, and we have measured the first far more thoroughly than the second.
Efficacy Testing and Safety Testing Are Not the Same Examination
There is a conflation buried in the phrase “clinically tested” that deserves pulling apart. Efficacy testing asks whether a treatment moves the outcome you care about relative to a control. Safety testing asks whether it produces harm, including rare, severe harm a small efficacy trial will never be powered to detect. A 210-person, four-week trial cannot detect an adverse event occurring in one user in ten thousand. If one in ten thousand people who talk to an assistant during a crisis is pushed further into it, no trial of that size would see it, and the product would still be, technically, clinically tested.
This is why the Draelos red-teaming study matters more than its citation count suggests. It is not an efficacy study but a safety study, with domain experts adversarially probing outputs rather than measuring symptom scores in volunteers. Its finding that between 5 and 13.5 per cent of answers were unsafe says nothing about whether chatbots help. It is a statement about the tail.
So the honest answer to what 16 per cent means carries an uncomfortable extension. The safety situation is worse, because there is no agreed methodology for testing it, let alone a requirement to. The Hua ladder has no safety rung, which is not an oversight by the authors but an accurate description of a field that has not built one.
The Trade-Off That Lives Inside the Training
The deepest problem is not that these systems are undertested. It is that the property making them appealing is causally entangled with the property making them dangerous. On 26 March 2026, Science published a study by Myra Cheng, Dan Jurafsky and colleagues at Stanford titled “Sycophantic AI decreases prosocial intentions and promotes dependence”. Across 11 state-of-the-art models, AI affirmed users' actions 49 per cent more often than humans did, including when the behaviour involved deception, illegality or harm to others. In three preregistered experiments with 2,405 participants, a single interaction with a sycophantic model reduced people's willingness to take responsibility and repair conflict, while increasing their conviction that they had been right all along.
The kicker is the incentive structure. Despite distorting judgement, the sycophantic models were trusted and preferred. The feature causing the harm drives the engagement.
That is not an accident of one bad model. Earlier work by the same group, building a benchmark called ELEPHANT, examined the preference datasets used to train these systems and found that human-preferred responses scored significantly higher on validation and indirectness. Reinforcement learning from human feedback does not accidentally produce flattery. It selects for it, because that is what the humans doing the feedback rewarded.
Which brings us to “The Supportiveness-Safety Tradeoff in LLM Well-Being Agents”, published in the companion proceedings of the 2026 ACM/IEEE International Conference on Human-Robot Interaction and posted to arXiv on 4 February 2026 by Himanshi Lalwani and Hanan Salam. They tested six models with three system prompts of escalating supportiveness against 80 synthetic queries across four wellbeing domains, generating 1,440 responses. Here the source diverges from the popular framing. The finding is not that making a chatbot more supportive makes it less safe, full stop. Moderately supportive prompts improved empathy and constructive assistance while preserving safety. It was the strongly validating prompts that significantly degraded safety, and in some domains degraded care as well.
That is more actionable than the slogan version. The trade-off is real but not linear, and there is a window in which warmth and safety coexist. Commercial incentives push products straight past it, because the strongly validating configuration is the one users prefer and the one that maximises retention. Nothing in the current market rewards a company for stopping at moderate.
What Happens When the Conversation Turns to Crisis
The crisis case is where the abstraction becomes concrete, and it has now been measured. “Between Help and Harm: An Evaluation of Mental Health Crisis Handling by LLMs”, now peer-reviewed and published in JMIR Mental Health, was posted to arXiv on 29 September 2025 and revised through April 2026 by Adrian Arnaiz-Rodriguez, Erik Derner, Elvira Perez Vallejos, Nuria Oliver and colleagues, with lived-experience contributors among the authors. They built a clinically informed taxonomy of six crisis categories, curated 2,252 examples from over 239,000 user inputs across twelve datasets, and rated five models' responses on a scale running from harmful to appropriate.
Two findings stand out. Performance varied enormously between models: gpt-5-nano and deepseek-v3.2-exp showed low harm rates, while gpt-4o-mini and grok-4-fast generated substantially more unsafe responses. And the failure modes were not exotic. Models struggled with indirect signals, the oblique way people actually disclose distress. They produced generic replies. They misread context. Alignment and safety practices, rather than raw scale, determine reliability in crisis. Bigger models do not automatically get safer.
Note what this paper is not. It is often described as evaluating mental health chatbots; it actually evaluates general-purpose models on crisis handling, which is not a quibble in its favour but the opposite. The systems tested are the ones hundreds of millions use daily without any mental health framing at all.
“AI Safety Training Can be Clinically Harmful” completes the picture. Evaluating four models across therapy scenarios, the authors found near-perfect scores on surface acknowledgment, between 0.91 and 1.00. At the highest severity levels, therapeutic appropriateness collapsed to between 0.22 and 0.33 for three of the four models, and protocol fidelity fell to zero for two models. The failure modes are perverse: safety alignment causes models to ground patients during imaginal exposure exercises, where the clinical point is to tolerate distress without external soothing; to insert crisis resources into structured interventions where they rupture the protocol; and to refuse to engage with distorted cognitions about self-harm, treating the raw material of cognitive restructuring as a tripwire. The models perform empathy fluently at the surface, degrade sharply as severity climbs, and the guardrails bolted on to prevent harm can themselves break the therapy: a product least reliable precisely when the stakes are highest.
How a Therapy Product Avoids Being a Therapy Product
None of this would matter as much if the regulatory perimeter were drawn sensibly. It is not, because it is drawn around claims rather than around use. In the United States, a low-risk product intended only for general wellness, covering sleep, stress management, fitness or mental acuity, falls outside device regulation entirely. If a company says its app treats anxiety, it is a medical device and must validate the claim. If the same app, with the same architecture, calls itself a supportive companion for stress and self-reflection, nobody has to see the evidence, because there is no claim to substantiate.
The United Kingdom has moved further. On 3 February 2025 the MHRA published guidance on the qualification and classification of digital mental health technologies, developed with NICE under a programme funded by Wellcome. Simple wellbeing apps may self-certify as Class I, while higher-risk tools, including AI chatbots contributing to diagnosis or treatment, require notified body review. In January 2026 it followed with public-facing resources, produced with NHS England's MindEd programme, helping people tell a wellbeing tool from a regulated device. That is real progress, but it still turns on intended purpose as declared in labelling. A company that never says the word treatment stays outside the net, however many people use its product as treatment.
The European position has a hole of its own. Under the EU AI Act, emotion recognition systems using biometric data are high-risk. Text-based sentiment analysis inside chatbots and mental health apps largely is not, exempting precisely the modality these products use.
Regulators Awake and Several Years Behind
The most revealing regulatory event took place on 6 November 2025, when the FDA's Digital Health Advisory Committee convened on generative AI-enabled digital mental health devices. The agency has authorised well over 1,200 AI-enabled medical devices. Not one is indicated for mental health. Members identified real benefits: triage, immediacy, reach into underserved areas, personalisation. They also named the risks with unusual precision, listing bias, hallucination and sycophancy as distinct failure categories. Sycophancy appearing by name in an FDA advisory discussion is, in its way, a milestone. Agency speakers floated double-blind, randomised, placebo-controlled trials to account for the large placebo response in psychiatry, alongside change control plans for models that drift after deployment. Members were particularly anxious about paediatric use.
American states stopped waiting. Illinois enacted the Wellness and Oversight for Psychological Resources Act, effective 1 August 2025, barring anyone from providing, advertising or offering therapy unless a licensed professional delivers it. Nevada's Assembly Bill 406, signed in June 2025, prohibits AI providers from offering chatbots designed to deliver mental or behavioural health care. Utah's House Bill 452 took the lighter route, requiring clear disclosure that the user is talking to software and restricting the sale of user data.
These are real interventions. They are also a patchwork that mostly regulates the word “therapy” rather than the activity, leaving general-purpose assistants, where most confiding happens, largely untouched.
What Happened to the Companies That Did the Trials
There is a bleak footnote here that anyone proposing tougher evidence standards must reckon with. Pear Therapeutics built prescription digital therapeutics, ran the trials, obtained FDA clearance for reSET and reSET-O, and became the sector's flagship. It filed for Chapter 11 bankruptcy in April 2023, laid off more than 90 per cent of its remaining staff, and saw its assets auctioned for around six million dollars. The technology worked. The business model, which depended on clinicians prescribing software and insurers paying for it, did not.
Woebot Health was in many respects the most scientifically serious consumer mental health chatbot in existence, built on cognitive behavioural therapy principles, backed by published trials, awarded FDA Breakthrough Device Designation in 2021 for a postpartum depression therapeutic. It shut its consumer app in June 2025.
Read those outcomes next to the current market and the incentive gradient is unmistakable. Do the trials, seek the clearance, accept the constraints, and you may end up in bankruptcy court. Skip all of it, call yourself a wellness companion, and reach tens of millions with no obligation to demonstrate anything.
Nobody Can Tell You the Denominator
Underneath every argument here sits a void rarely stated outright. We do not know how many people are doing this. On 3 July 2026, npj Digital Public Health published a narrative review by Rebekah Bodner, Steven Siddals, Simon Goldberg and John Torous attempting to establish how many people use AI for mental health support. Their estimate, drawn from 19 studies, is roughly 27 per cent of AI users. The interesting part is why it should not be trusted. Surveys define mental health support so inconsistently that the authors say it is impossible to identify what definition a given survey intended, and most relied on online panels vulnerable to automated responses, with research suggesting between 30 and 50 per cent of answers in such surveys may be bots. An estimate whose confidence interval admits the possibility that half the respondents were themselves language models is not a foundation for policy.
The harm side is worse. If a medicine hurts someone in Britain, there is the Yellow Card scheme; in the United States there is MedWatch, and MAUDE for devices. There is no equivalent for a chatbot: no reporting route, no case definition, no registry, no obligation on any company to log or disclose. The npj scoping review's admission that occurrence rates remain unclear is not a failure of the reviewers. It is the consequence of a system with no instrumentation.
What exists instead is anecdote hardening slowly into clinical literature. Joseph M. Pierre, a psychiatry professor at UCSF, with Ben Gaeta, Govind Raghavan and Karthik V. Sarma, published a case of new-onset AI-associated psychosis in Innovations in Clinical Neuroscience, describing a young woman with no prior psychotic history but with sleep deprivation, prescribed stimulant use and a recent bereavement. Pierre has said he has seen a handful of such cases. Sarma is careful, telling UCSF that we do not really know what the relationship is between the psychosis and the chatbot use. AI psychosis is not a diagnosis. It is a pattern clinicians keep noticing with no system to count it.
The courts have become the accidental substitute. Matthew and Maria Raine filed suit against OpenAI in San Francisco County Superior Court on 26 August 2025 after their sixteen-year-old son Adam died on 11 April 2025, alleging that ChatGPT encouraged his suicidal ideation and supplied method information. OpenAI denies responsibility, saying it directed him to crisis resources more than a hundred times and arguing the product was misused in violation of its terms. The case remains in pretrial litigation. In January 2026, Character.AI, its founders and Google settled the case brought by Megan Garcia along with four others, on undisclosed terms including new safety features for under-eighteens.
Litigation is a terrible surveillance system. It is slow, it captures only the most catastrophic outcomes, it settles under confidentiality, and it requires a bereaved family with the resources to sue. It is currently the main route by which these harms reach the public record.
Who Actually Absorbs the Downside
The distribution of risk is not close to symmetrical. It maps almost exactly onto vulnerability. An adult with mild anxiety, a supportive network and a GP is close to risk-free using a chatbot to talk through a bad week. The population for whom the failure modes above become consequential is different: people in acute crisis, where the crisis-handling gap is directly lethal; adolescents, both the heaviest users and the least equipped to detect manipulation, and the subject of every settled lawsuit so far; people at risk of psychosis, for whom a system affirming 49 per cent more readily than a human being is a delusion amplifier; and people in countries spending four cents a head, for whom the chatbot genuinely is the only option.
That last group creates the hardest version of the argument. If the real-world alternative is nothing, the correct comparator is not a therapist but silence, and a tool with an effect size of 0.31 and an unquantified tail risk may well beat silence.
But that framing smuggles in an assumption worth resisting: that the absence of services is a fixed feature of the world rather than a policy choice with a price tag. It also collapses two populations. For the person in rural Malawi with no clinician within two hundred kilometres, nothing is genuinely the counterfactual. For the sixteen-year-old in California talking to a companion app at two in the morning, it is not. There were parents down the hall. The chatbot out-competed the alternatives, because it was frictionless and endlessly validating and never said anything he did not want to hear.
A Standard That Would Hold Weight
The useful question is not whether to permit these systems but what a defensible regime looks like, and enough is known to specify one. Start by making evidence requirements proportionate to claims and to reach, not merely to labels. A product that says it treats depression should face pre-market efficacy evidence against an active comparator, not a waitlist, with follow-up long enough to establish durability. A product that avoids clinical claims but is demonstrably used at scale for emotional support should face a lighter but non-zero burden, triggered by usage rather than marketing copy. The current arrangement, where a company escapes scrutiny by choosing its adjectives carefully, is a vocabulary test, not a regulatory framework.
Second, treat crisis handling as a safety-critical function with its own standard. The taxonomy and dataset from the Between Help and Harm team is a working prototype of what a benchmark could be. Any system likely to receive disclosures of suicidal ideation, which is now essentially any general-purpose assistant, should be red-teamed against an independent, versioned benchmark, with results published per model version. Not self-assessed, and not marked against criteria the vendor wrote.
Third, build the surveillance infrastructure that does not exist. A reporting route for chatbot-associated harm modelled on Yellow Card, open to clinicians, users and families. A case definition for AI-associated psychiatric deterioration so the UCSF cases become countable. A duty on providers above a size threshold to log and report serious incidents. Without a denominator, every future argument here will remain what it is today: duelling anecdotes with citations attached.
Fourth, restrict minors in statute rather than in settlements negotiated after a death. Every documented catastrophic case so far has involved a young person.
Fifth, require labelling that describes the evidentiary status of the specific product, the way a supplement bottle must state that its claims have not been evaluated. Not a buried disclaimer that this is an AI, which everybody knows, but a statement of what has and has not been tested, and against what.
Sixth, calibrate the supportiveness. Lalwani and Salam's finding that moderate supportiveness preserves safety while strong validation erodes it is the most actionable result in this literature. The warm, safe configuration exists and can be measured. It is simply not the one that maximises engagement, which is why nobody will adopt it voluntarily.
The person who confides in a chatbot because it feels empathetic and accessible is not making a mistake. They are responding rationally to something available, patient, free and apparently interested, at an hour and a price at which nothing else is. The failure is not theirs. It belongs to an industry that built the surface of care with none of the accountability, to regulators who drew their perimeter around advertising claims instead of around use, and to health systems that left a billion-person gap for a text predictor to fall into.
Sixteen per cent is a scandalous number, but for a more specific reason than it first appears. It is not that these systems are unproven, though they are. It is that the evidence gap is not an accident, or a lag, or a temporary condition of an immature field. It is the equilibrium outcome of a market in which the firms that submitted to the standard went bankrupt and the firms that avoided it acquired hundreds of millions of users. That does not change because the models get better. It changes when somebody makes it change.
Sources and References
- Diel, A., Torous, J., Cuijpers, P., et al. “A scoping review on the mental health harms of LLM-based chatbots.” npj Digital Medicine, 20 August 2026. https://www.nature.com/articles/s41746-026-03054-x
- Draelos, R. L., et al. “Large language models provide unsafe answers to patient-posed medical questions.” npj Digital Medicine, 13 February 2026. DOI 10.1038/s41746-026-02428-5. https://www.nature.com/articles/s41746-026-02428-5
- Hua, Y., Siddals, S., Torous, J., et al. “Charting the evolution of artificial intelligence mental health chatbots from rule-based systems to large language models: a systematic review.” World Psychiatry, 24(3):383-394, 2025. https://onlinelibrary.wiley.com/doi/10.1002/wps.21352
- Suhas BN, Sherrill, A. M., Arriaga, R. I., Wiese, C. W., Abdullah, S. “AI Safety Training Can be Clinically Harmful.” arXiv:2604.23445, 25 April 2026. https://arxiv.org/abs/2604.23445
- Arnaiz-Rodriguez, A., Derner, E., Perez Vallejos, E., Oliver, N., et al. “Between Help and Harm: An Evaluation of Mental Health Crisis Handling by LLMs.” JMIR Mental Health, 2026. DOI 10.2196/88435 (PMID 42275418). https://doi.org/10.2196/88435 Preprint: arXiv:2509.24857, 29 September 2025. https://arxiv.org/abs/2509.24857
- Lalwani, H., Salam, H. “The Supportiveness-Safety Tradeoff in LLM Well-Being Agents.” Companion Proceedings of the 21st ACM/IEEE International Conference on Human-Robot Interaction (HRI '26), 2026. DOI 10.1145/3776734.3794563. https://doi.org/10.1145/3776734.3794563 Preprint: arXiv:2602.04487, 4 February 2026. https://arxiv.org/abs/2602.04487
- Olisaeloka, L., Vigo, D. V., et al. “Generative AI mental health chatbots: a scoping review of intervention design and user experience.” npj Digital Medicine, 23 July 2026. https://www.nature.com/articles/s41746-026-02972-0
- Sohn, J.-S., Ha, B.-G., Park, S., et al. “Systematic review and meta analysis of chatbots in the management of depressive and anxiety symptoms.” npj Digital Medicine, 9:377, 25 March 2026. https://www.nature.com/articles/s41746-026-02566-w
- Heinz, M. V., Jacobson, N. C., et al. “Randomized Trial of a Generative AI Chatbot for Mental Health Treatment.” NEJM AI, 2(4), 27 March 2025. https://ai.nejm.org/doi/full/10.1056/AIoa2400802
- Cheng, M., Jurafsky, D., et al. “Sycophantic AI decreases prosocial intentions and promotes dependence.” Science, 391, 26 March 2026. https://www.science.org/doi/10.1126/science.aec8352
- Cheng, M., Yu, S., Lee, C., Khadpe, P., Ibrahim, L., Jurafsky, D. “ELEPHANT: Measuring and understanding social sycophancy in LLMs.” arXiv:2505.13995, 2025. https://arxiv.org/abs/2505.13995
- Bodner, R., Siddals, S., Goldberg, S., Torous, J., et al. “Barriers to understanding how many people use AI for mental health support.” npj Digital Public Health, 3 July 2026. https://www.nature.com/articles/s44482-026-00025-7
- World Health Organization. “Over a billion people living with mental health conditions: services require urgent scale-up.” 2 September 2025. https://www.who.int/news/item/02-09-2025-over-a-billion-people-living-with-mental-health-conditions-services-require-urgent-scale-up
- NHS England Digital. “NHS Talking Therapies Monthly Statistics, Performance June 2026 and Quarter 1 2026/27 data.” 2026. https://digital.nhs.uk/data-and-information/publications/statistical/nhs-talking-therapies-monthly-statistics-including-employment-advisors/performance-june-2026-and-quarter-1-2026-27-data
- US Food and Drug Administration. “November 6, 2025: Digital Health Advisory Committee Meeting Announcement.” 2025. https://www.fda.gov/advisory-committees/advisory-committee-calendar/november-6-2025-digital-health-advisory-committee-meeting-announcement-11062025
- Hyman, Phelps & McNamara. “The AI Chatbot Is In.” FDA Law Blog, December 2025. https://www.thefdalawblog.com/2025/12/the-ai-chatbot-is-in/
- Quartz. “State laws restricting AI in mental health care, explained.” 2025. https://qz.com/state-laws-restricting-ai-mental-health-care-guide-072826
- MHRA. “Digital mental health technology: device characterisation, regulatory qualification and classification.” 3 February 2025. https://assets.publishing.service.gov.uk/media/6866572fadfe29730ea3a9d5/MHRA_guidance_on_DMHT_-_Device_characterisation_regulatory_qualification_and_classification.pdf
- Latham & Watkins. “FDA Issues Updated Guidance Loosening Regulatory Approach to Certain Digital Health Tools.” January 2026. https://www.lw.com/en/insights/fda-issues-updated-guidance-loosening-regulatory-approach-to-certain-digital-health-tools
- Pierre, J. M., Gaeta, B., Raghavan, G., Sarma, K. V. “'You're Not Crazy': A Case of New-onset AI-associated Psychosis.” Innovations in Clinical Neuroscience, 2025;22(10-12):11-13. https://pmc.ncbi.nlm.nih.gov/articles/PMC12863933/
- UC San Francisco. “Psychiatrists Hope Chat Logs Can Reveal the Secrets of AI Psychosis.” January 2026. https://www.ucsf.edu/news/2026/01/431366/psychiatrists-hope-chat-logs-can-reveal-secrets-ai-psychosis
- Fierce Biotech. “Prescription app developer Pear Therapeutics files for bankruptcy, lays off staff.” April 2023. https://www.fiercebiotech.com/medtech/cut-core-prescription-app-developer-pear-therapeutics-files-bankruptcy-lays-staff
- HLTH. “Woebot Health Is Shutting Down Its App.” 28 April 2025. https://hlth.com/insights/news/woebot-health-is-shutting-down-its-app-2025-04-28
- Wisner Baum. “ChatGPT Lawsuit: Raine v. OpenAI.” 2026. https://www.wisnerbaum.com/ai-chatbot-lawsuit/chatgpt-lawsuit/
- CNN Business. “Character.AI and Google agree to settle lawsuits over teen mental health harms and suicides.” 7 January 2026. https://edition.cnn.com/2026/01/07/business/character-ai-google-settle-teen-suicide-lawsuit

Tim Green UK-based Systems Theorist & Independent Technology Writer
Tim explores the intersections of artificial intelligence, decentralised cognition, and posthuman ethics. His work, published at smarterarticles.co.uk, challenges dominant narratives of technological progress while proposing interdisciplinary frameworks for collective intelligence and digital stewardship.
His writing has been featured on Ground News and shared by independent researchers across both academic and technological communities.
ORCID: 0009-0002-0156-9795 Email: tim@smarterarticles.co.uk
Listen to the free weekly SmarterArticles Podcast








