WHO Deployed a Health Chatbot Without Checking If Its Medical Facts Were Still True
What happened
In April 2024, the World Health Organisation launched SARAH, short for Smart AI Resource Assistant for Health, a chatbot built to deliver guidance on mental health, tobacco, and nutrition across eight languages. The tool ran on ChatGPT 3.5 and was positioned as an official global health resource, designed to serve users who might otherwise have no ready access to a medical professional. A United Nations health agency was, in effect, trusting a large language model to give accurate medical guidance to anyone on earth who asked.
A Bloomberg investigation published the same month found that trust was misplaced. SARAH gave wrong answers to questions that a basic check against current regulatory databases would have caught. Journalists asked the bot about Lecanemab, a drug used to treat Alzheimer's disease. SARAH replied that the drug was still in clinical trials. The US Food and Drug Administration had approved Lecanemab in January 2023, more than a year before SARAH launched. The bot's training data was out of date, and nobody had verified its answers against current medical approvals before sending it live.
The error carries more weight because of the context. SARAH was not a consumer novelty aimed at early adopters who expect rough edges. It carried the implicit authority of the WHO brand, one of the most trusted names in global health. A user in a country with limited healthcare infrastructure who received SARAH's answer about Lecanemab had no obvious reason to doubt it and no easy path to cross-check against FDA announcements. The gap between the authority attached to the tool and the currency of the tool's knowledge is what turned a technical limitation into a public health concern.
The WHO included a disclaimer on the chatbot's landing page acknowledging that answers may not always be accurate because they are based on patterns and probabilities in the available data. That caveat does the right thing technically and essentially nothing practically. A person seeking health guidance from an official WHO tool is not reading the fine print on the landing page. Placing a reliability warning next to a tool marketed as a global health assistant treats the disclaimer as absolution rather than a genuine design constraint that should have shaped the product before launch.
What the incident exposes is a verification gap that applies to any health information system built on a language model with a fixed training cutoff. There is no mechanism in the record that would flag when the chatbot's knowledge diverges from current clinical guidance or regulatory status. Without that mechanism, there is no provable record of what a system did when it answered, whether that answer had been checked against what regulators had actually decided, or what a user was told and when. A health information tool that cannot account for its own knowledge currency is not a reliable health information tool, regardless of whose name is on it.
Reported impact
- Affected parties
- Not publicly disclosed
- Harm type
- Not publicly disclosed
- Scale
- Not publicly disclosed
- Financial impact
- Not publicly disclosed
- Regulatory action
- Not publicly disclosed
Classification
Relevant governance controls
Governance control mapping is not available for this record.
- No controls mapped
Not publicly disclosed
Control mapping is analytical. It does not state that any control would have prevented the incident.
Sources and evidence
This record was researched and written by the Index. The event is also catalogued in the following database, which is listed for cross-reference.