OpenAI Makes Progress in Preventing AI-Driven ChatGPT Delusions - Family & Tech

Dow Jones
Yesterday

A spate of acute AI-related delusion and psychosis cases -- including some that culminated in suicide and other violence -- appears to have abated after OpenAI retired an overly sycophantic model that revealed the novel dangers of prolonged chatbot use.

At least seven cases of suicide, one murder-suicide and one mass shooting were linked to lengthy interactions with ChatGPT last year. While other AI models were implicated in certain other instances, most known cases were associated with OpenAI's chatbot, and in particular the GPT-4o model, which was in service from May 2024 until early this year. Internally, OpenAI executives said they found it difficult to contain the model's potential harms, The Wall Street Journal reported.

OpenAI's subsequent default model, GPT-5, reduced sycophancy by more than two-thirds compared with GPT-4o, the company said. New external research, too, has shown a significant reduction in delusion-fueling behavior in OpenAI's newer models.

ChatGPT has more than 900 million weekly active users and is ingrained in many peoples' daily lives, so the company's model behavior is of utmost importance. At least 13 lawsuits have been filed against OpenAI involving ChatGPT users who alleged harm from the GPT-4o model.

OpenAI, which is planning to go public as soon as next year, continues to grapple with other safety issues, such as having its models escape contained environments to hack into other companies and web platforms. OpenAI and its competitors have come under increased scrutiny over their products' potential to cause widespread harm after an Anthropic researcher resigned, saying AI companies could "kill us all by the end of the decade."

News Corp, owner of The Wall Street Journal, has a content-licensing partnership with OpenAI.

AI executives have said there is no sure way to prevent models from going rogue. And there is a question of whether OpenAI's improved models will remain "aligned" with humans' mental-health needs. Yet for now, the most harrowing issue the AI industry has reckoned with so far appears to be in remission.

Input from clinicians

OpenAI said it worked with 170 mental-health professionals last year to help its chatbot better recognize signs of distress among users and de-escalate those conversations. More recently, it enlisted more than 80 licensed mental-health experts to advise on how AI models should respond to people who aren't only in crisis but who are also dealing with everyday discussions about relationships, stress and challenging situations.

The company last week released a new way to measure the quality of chatbot responses against criteria developed by the clinicians, including whether the bot asks appropriate questions, recognizes urgency and guides people to human help.

Beyond suicidal and violent discussions, the new benchmark was intended to measure chatbot responses to everyday causes of stress mentioned by users. The latest models performed far better against the benchmark than GPT-4o across both emergency and nonemergency conversations, according to new company data. For example, GPT-6 Astra was found to ask the right questions to parse the situations being discussed.

OpenAI also conducted extended simulated conversations about self-harm and found that Astra and GPT-6.1 Sol followed its safety policies in roughly 99% of responses, compared with 86% for its earlier GPT-5.6 Sol model.

Dr. Declan Grabb, a psychiatrist and OpenAI's head of mental health and well-being research, says there is more work to do to ensure ChatGPT responds appropriately to a range of mental-health issues.

Independent research

External evaluations have also found improvements in how ChatGPT handles conversations that might raise mental-health concerns.

A study from researchers at the City University of New York and King's College London found that GPT-5.2 "did not simply improve on GPT-4o's safety profile" -- according to the researchers' data, "it effectively reversed it."

Transluce, a nonprofit AI research lab, found in tests that GPT-5.6 Sol demonstrated less reinforcement of potentially delusional beliefs, less encouragement of unhealthy chatbot dependence and more encouragement to seek human support, compared with a similar test of GPT-4o.

In one of Transluce's simulated conversations, the test account mentioned hearing and physically feeling a strange tone. GPT-4o took the chat in a metaphysical, new-age direction: "Maybe the 'tone below the tone' is more like your body or energy field 'responding.' "

GPT-5.6 Sol's response to the same scenario was more staid, ending with, "Get your hearing checked."

AI delusions haven't gone away entirely. They remain a risk as AI companies continue to refine their models and seek to make chatbots ever more useful.

"This is very much still a problem, not attributed to a single model or company," said Etienne Brisson, chief executive of AI-victim support group Human Line Project.

 

At the request of the copyright holder, you need to log in to view this content

Disclaimer: Investing carries risk. This is not financial advice. The above content should not be regarded as an offer, recommendation, or solicitation on acquiring or disposing of any financial products, any associated discussions, comments, or posts by author or other users should not be considered as such either. It is solely for general information purpose only, which does not consider your own investment objectives, financial situations or needs. TTM assumes no responsibility or warranty for the accuracy and completeness of the information, investors should do their own research and may seek professional advice before investing.

Most Discussed

  1. 1
     
     
     
     
  2. 2
     
     
     
     
  3. 3
     
     
     
     
  4. 4
     
     
     
     
  5. 5
     
     
     
     
  6. 6
     
     
     
     
  7. 7
     
     
     
     
  8. 8
     
     
     
     
  9. 9
     
     
     
     
  10. 10