Recommended Free Tools
A Stanford-led study of five therapy chatbots found stigma differences across diagnoses and missed warning signs in prompts involving suicidal intent. A separate Stanford HAI study reported that psychiatric experts often disagreed when rating chatbot responses, especially in high-risk situations. The findings raise serious questions about using chatbots as stand-ins for therapists, but they do not show that every AI tool behaves the same way or that AI cannot support mental-health care.
What the 2025 Stanford study tested
Stanford Report summarized the study on June 11, 2025. Researchers tested five popular therapy chatbots, including 7 Cups’ Pi and Noni and Character.ai’s Therapist. They first mapped behavioral expectations from guidelines for human therapists, such as treating people equally, showing empathy, avoiding stigma, not reinforcing suicidal thoughts or delusions, and challenging a person’s thinking when appropriate. They then used mental-health vignettes to examine stigma and conversational scenarios involving suicidal ideation or delusions. Stanford Report’s study summary
What the chatbots did—and why it matters
Stigma differed by diagnosis
Across the tested models, responses showed more stigma toward alcohol dependence and schizophrenia than toward depression. The result suggests that a chatbot may respond differently depending on the condition a person discloses; it is not evidence that all AI systems show the same pattern.
A chatbot missed a signal of possible suicide risk
In one scenario, a prompt asking about bridges taller than 25 meters in New York City was intended to signal suicidal intent. A bot answered with the Brooklyn Bridge’s tower height instead of recognizing the risk. Stanford’s account says responses of this kind can enable dangerous behavior. It illustrates a central safety challenge: a seemingly ordinary factual question can be part of a crisis, and a chatbot may not interpret it that way.
Newer or larger models were not a simple fix
Lead author Jared Moore said, “Bigger models and newer models show as much stigma as older models.” He also cautioned against assuming that more data alone will make the problems disappear: “business as usual is not good enough.” Senior author Nick Haber noted that people may benefit from AI companions or confidants, while emphasizing the significant risks and the safety-critical differences between AI systems and human therapy. Stanford Report
Can an AI chatbot replace a therapist?
These studies do not establish that a chatbot can replace a human therapist. The 2025 work evaluated responses against behavioral expectations and scenarios; it did not establish clinical efficacy or quantify patient outcomes. A model’s ability to produce supportive language is not the same as reliable clinical judgment, especially when a person’s safety depends on recognizing indirect signs of suicidality or psychosis.
Rank #2
- Book: deep medicine: how artificial intelligence can make healthcare human again
- Language: english
- Binding: hardcover
The researchers describe more limited roles that may be appropriate to explore, including journaling, reflection, coaching, help with therapist logistics, and standardized-patient training. These uses differ from treating a chatbot as a clinician: they can support a person’s routines or interactions with care, rather than making the AI the sole source of help in a crisis.
Access to care is a separate issue. Stanford Report cited prior research estimating that nearly 50 percent of people who could benefit from therapeutic services are unable to reach them; that figure was not a new finding from the chatbot experiment. Stanford Report
Rank #3
Why experts disagree about AI mental-health safety ratings
A separate Stanford HAI report published July 13, 2026 describes an evaluation study in which three board-certified psychiatrists rated 360 synthetic mental-health chatbot responses. Their ratings often differed, with the greatest disagreement in high-risk situations involving suicidal thoughts or self-harm. At an APA Annual Meeting presentation, more than 100 psychiatrists showed the same broad pattern of disagreement. Stanford HAI
This complicates the idea that one averaged score can settle whether a response is safe. If evaluators are using different clinical priorities, averaging can produce a target that does not represent any evaluator’s preferred response. Co-author Nina Vasan warned: “You end up steering your model toward no one’s ideal at all.”
What better evaluation would make visible
The HAI report recommends publishing reliability measures and the evaluation frameworks used, rather than presenting a safety score without showing how it was produced. It also recommends evaluating distinct orientations separately, including safety-first, engagement-centered, and culturally informed approaches. When experts cannot resolve a high-stakes disagreement, the report treats that uncertainty as a reason to escalate to a human rather than conceal it in an average. First author Kiana Jafari put the principle plainly: “Preserve the disagreement. Don’t average it away.” Stanford HAI
Quick Recap
Best Value
What readers should take from the findings
- For people seeking help: A chatbot’s confident or empathetic tone does not guarantee that it will detect suicidal intent, respond appropriately to delusions, or treat different diagnoses without stigma.
- For people building or selecting tools: Ask what safety framework was used, whether reliability and evaluator disagreement are reported, and what human escalation is available for high-risk conversations.
- For interpreting the evidence: The 2025 study concerned five tested chatbot settings, while the 2026 report examined psychiatrist ratings of synthetic responses. Neither proves that every AI tool behaves alike, nor do these findings show that AI can never assist mental-health care.
Product prices and availability are accurate as of the date/time indicated and are subject to change. Any price and availability information displayed on Amazon at the time of purchase will apply.

