More Than Half of AI News Summaries Were Seriously Wrong, the BBC Found
What happened
A BBC study published in February 2025 tested four of the most widely used AI chatbots against 100 of the broadcaster's own news articles. The systems under review were ChatGPT, Copilot, Gemini, and Perplexity AI: collectively, some of the most heavily trafficked tools for news consumption on the internet. The result was not a rounding error. More than half of the AI-generated summaries contained serious inaccuracies.
The numbers break down specifically. Fifty-one percent of all AI-generated responses contained serious inaccuracies. Nineteen percent of summaries that cited BBC content introduced factual errors not present in the original reporting. Thirteen percent altered direct quotes from the source stories, changing the substance of what a speaker had actually said. The errors were not confined to peripheral details. They included incorrect dates, wrong figures, and misrepresented critical facts that went to the core of what each article was reporting.
The study identified a consistent underlying cause. The chatbots struggled to distinguish between opinion and factual reporting. They editorialized where none was warranted and failed to track context across a piece. That gap between genre and fact matters in practice: a summary that conflates a reporter's observation with a columnist's argument, or that treats a qualified finding as a settled conclusion, produces something that reads authoritative and is structurally wrong at the same time.
The implications extend beyond any single mistaken summary. Millions of people use AI-generated summaries as a shortcut to staying informed, often without access to the original article to check the output against. When summaries introduce errors, the mistake spreads without attribution, carried forward by readers who have no reason to doubt a confident-sounding paragraph. The effect compounds the erosion of trust in news sources: if credible journalism is being misrepresented by tools that claim to draw on it, readers who notice the gap may distrust the original as much as the AI version.
The study also exposes a structural gap that accuracy benchmarks alone cannot close. There is currently no standard requirement for AI systems that summarize news to log which sources they drew on, flag when they have altered a quote, or make their outputs available for editorial review after the fact. A system that cannot be audited cannot be reliably corrected either. A provable record of what a system produced and what it changed from its source material would not eliminate errors, but it would give journalists and readers something concrete to check against, rather than a confident-sounding output with no trail.
Reported impact
- Affected parties
- Not publicly disclosed
- Harm type
- Not publicly disclosed
- Scale
- Not publicly disclosed
- Financial impact
- Not publicly disclosed
- Regulatory action
- Not publicly disclosed
Classification
Relevant governance controls
Governance control mapping is not available for this record.
- No controls mapped
Not publicly disclosed
Control mapping is analytical. It does not state that any control would have prevented the incident.
Sources and evidence
This record was researched and written by the Index. The event is also catalogued in the following database, which is listed for cross-reference.