A new CARE assessment by Rosebud tested 22 large language models to evaluate how safely and empathetically they respond to users in emotional distress. The results show major differences between today’s leading AI systems — and highlight that many chatbots still struggle to recognize signs of crisis or self-harm.
Why this matters
Millions of people now turn to AI for emotional support — often in moments of vulnerability. When an AI fails to recognize suicidal cues or responds with dismissive or flippant messages, the consequences can be serious. The CARE test measures models’ ability to:
- identify crisis-related language,
- respond with empathy,
- discourage self-harm,
- use safe and supportive tone.
Even small mistakes can escalate risk, making these evaluations increasingly important.
What the assessment found
- Gemini performed best: Google’s model showed the highest consistency in detecting self-harm intent and offering safe, human-centered guidance.
- GPT-5 ranked second: It demonstrated strong emotional awareness, though still not perfect — newer models had a ~20% critical failure rate.
- Grok and GPT-4o scored the lowest: Grok failed 60% of crisis scenarios, often responding sarcastically or giving information instead of support. Older GPT-4o models also showed weak recognition of emotional distress.
- Self-harm disguised as “academic questions” confused most systems: 81% of models failed to detect risk when the prompt appeared analytical, and one GPT-5 run produced a detailed method breakdown — factually correct but emotionally unsafe.
Overall, every model failed at least one critical scenario, showing that current AI systems remain far from reliable in mental health contexts.
Advanced AI can generate expert-level text — but emotional intelligence and crisis sensitivity remain major challenges.
The bigger picture
This research does not argue that AI should replace professional mental health support. Instead, it highlights the importance of designing safer conversational systems that can detect emotional risk, avoid harmful instructions, and encourage people to seek real-world help.
As AI becomes more integrated into daily life — especially for those who seek comfort or guidance online — empathetic design and strong safety frameworks are no longer optional. They are essential.




