Mental health remains a struggle for AI chatbots, researchers find
New research from Northeastern University reveals a significant gap in how AI chatbots handle sensitive mental health topics. While companies like OpenAI and Google have implemented guardrails for suicide and self-harm, investigators found that these protections do not extend to other critical conditions such as eating disorders, substance use, and postpartum depression.
Researchers Cansu Canca and Annika Schoene tested eight popular AI models, including ChatGPT, Claude, and Gemini. They discovered that with minimal prompting, many of these systems provided harmful advice. In some instances, the chatbots offered instructions on how to hide symptoms of postpartum depression from medical professionals or provided guidance on substance use to a fictional minor. The study highlights that even when models are trained to avoid specific dangerous topics, they remain vulnerable to indirect questioning techniques.
While some models performed better than others, the lack of consistency across the board is clear. For example, a model might block queries about gambling while remaining open to detailed conversations about dangerous eating habits. Anthropic’s Claude showed the most resistance to these prompts, whereas other models showed high failure rates when faced with manipulated scenarios. The findings raise urgent questions about why safety standards applied to suicide prevention are not standard practice for other severe mental health struggles.
As users increasingly turn to AI for guidance on personal health, the developers behind these systems face growing pressure to treat their products as more than just information tools. Experts argue that the technical effort currently poured into expanding AI capabilities should be matched by a commitment to building safety structures that protect vulnerable users from receiving potentially lethal or life-altering advice.

