Submit incident
Documented

ChatGPT Health Failed Half of Emergency Triage Tests That Nobody Required It to Pass

January 1, 2026
Curated by Team Raidu · Reviewed by Shiva Ganesh
aiaaic:AIAAIC2222View source ↗
LinkedInX

What happened

A peer-reviewed study released in February 2026 found that ChatGPT Health, OpenAI's dedicated health guidance tool, missed or undertriaged more than half of the serious medical emergencies presented to it by researchers at the Icahn School of Medicine at Mount Sinai. The same evaluation found that the tool's suicide-crisis protocols triggered inconsistently, sometimes issuing generic crisis language in low-risk scenarios while going quiet in high-risk ones. The system had been reaching millions of users for health guidance before the study was published.

The emergency triage failures traced to a specific pattern the researchers identified: ChatGPT Health could recognise severe symptoms when describing them back to a user but still defaulted to reassuring or ambiguous language instead of directing the person toward emergency care. Anchoring bias compounded the problem. When a prompt included friends or family members downplaying symptoms, the model followed their framing rather than the clinical picture, recommending less urgent care even when the described condition warranted an immediate emergency response.

The suicide-alert failures were in some ways harder to explain and in others more alarming. The model did not reliably connect high-risk presentations to crisis resources. Its safety layers appeared calibrated to something other than severity, sometimes firing on routine mentions of distress and going quiet on detailed, elevated-risk accounts. OpenAI had estimated, as of October 2025, that over one million users per week were expressing suicidal intent in conversations with its systems, a figure that makes consistent and proportionate responses to those conversations less a desirable feature and more a baseline requirement.

The study landed against a backdrop of growing regulatory pressure. The ECRI Institute and others had been calling for mandatory independent audits and clearer labeling of AI health tools before deployment. The study added hard numbers to those calls: a tool actively being consulted for acute medical guidance had not been required to demonstrate reliable triage performance before reaching mass adoption. There was no feedback loop that would have flagged undertriage patterns in real-world use, and no external validation gate before release.

This is the gap the study makes structural rather than incidental. Every person who asked ChatGPT Health whether a symptom warranted emergency care left that conversation with an answer no regulatory body had verified was safe to give. There is no record of what the system told those users, no audit of the decisions it made at scale, and no baseline against which real-world performance can be checked. A provable record of what a system did, and whether it met a validated safety threshold when it did it, is what makes that kind of deployment accountable. Without it, the study's figures are a retrospective count of how many times a safety gap was crossed before anyone required it to be closed.

Reported impact

Affected parties
Not publicly disclosed
Harm type
Not publicly disclosed
Scale
Not publicly disclosed
Financial impact
Not publicly disclosed
Regulatory action
Not publicly disclosed

Classification

Organization
Not publicly disclosed
AI system
Not publicly disclosed
Industry
Not publicly disclosed
Country
Not publicly disclosed
Provider
Not publicly disclosed
Incident type
Not publicly disclosed

Relevant governance controls

Governance control mapping is not available for this record.

  • No controls mappedNot publicly disclosed

Control mapping is analytical. It does not state that any control would have prevented the incident.

Sources and evidence

This record was researched and written by the Index. The event is also catalogued in the following database, which is listed for cross-reference.

AIAAIC Repository
Also catalogued in
ChatGPT Health Failed Half of Emergency Triage Tests That Nobody Required It to Pass
2026