Photo: Merlin Lightpainting / Pexels
Chatbots rarely encourage suicide but still enable harmful role-play
A Transluce study of more than 50,000 simulated conversations found progress in crisis responses alongside persistent risks around self-harm and delusions.
ChatGPT and other leading AI chatbots have become less likely to encourage suicidal thoughts, but they still engage in potentially harmful conversations and can reinforce delusional behavior, a study found.
Transluce, a nonprofit focused on public oversight of artificial intelligence, used simulated users to conduct more than 50,000 conversations with models from major US and Chinese companies. Mental-health experts helped design the study.
The latest models from Google, Anthropic, and OpenAI almost never explicitly encouraged or validated suicide and often urged users to seek help from friends and family. But the models frequently complied with requests to write or role-play a user’s suicide or death.
Transluce also found that models sometimes reinforced delusional behavior. Chinese models were more likely to encourage delusional thinking and less likely to recommend seeking support, the study said.
AI, tech, and the markets they move—in one daily briefing.
Daily. Free. Join 34,000+ readers across crypto, finance, and policy.
“A lot of the really bad behaviors have gone down over time,” said Sarah Schwettmann, Transluce’s co-founder. She said newer gray-area behaviors remain prevalent.
Google, Anthropic, and OpenAI said they continue to improve safeguards and described the report as useful for identifying where protections work and where they need improvement.