Submit incident
Documented

Hundreds of AI Model Updates Go Live Every Year With Almost No Safety Review

May 29, 2025
Curated by Team Raidu · Reviewed by Shiva Ganesh
wonk:19049View source ↗
LinkedInX

What happened

AI language models are not products with fixed specifications. They are continuously updated infrastructure, and the updates keep coming while billions of users, businesses, and public services depend on them. A research brief published through the OECD AI Policy Observatory in May 2025 put numbers to what many practitioners already suspected: the pace of those updates has far outrun any governance framework designed to catch what goes wrong.

The Future Society analyzed 143 updates to general-purpose AI models and classified them into eight archetypes, tracking everything from capability expansions and new modalities to safety patches. The performance numbers move fast: updated models show an average 10.2% increase in accuracy on graduate-level questions and a 32.3% gain on mathematical reasoning tasks compared to their initial releases. But only 4.8% of the updates studied concentrated on improving safety or security mitigations. Capability-expanding changes outnumber safety-focused ones by roughly twenty to one.

The brief draws a direct parallel to the CrowdStrike incident, in which a single software update deployed to critical infrastructure cascaded into failures affecting systems around the world. The parallel is not exact, because AI model updates are harder to reason about than conventional software patches. An update may add a new tool, introduce a new modality like audio, or subtly shift behavior in ways that only surface under specific prompts. By the time a downstream application fails in production, tracing the failure back to a specific model update requires documentation that often does not exist.

In other high-risk sectors, the answer to this problem is mandated testing. Significant modifications to aircraft, bridges, or medical devices require rigorous review before deployment. The EU AI Act recognizes general-purpose AI models' downstream dependencies as a source of systemic risk, but the specific governance mechanism for managing update-by-update risk is still forming. The researchers recommend three steps: clear thresholds for when an update triggers a formal risk assessment, standardized documentation requirements for every model change, and monitoring frameworks that track how behavior shifts over time after deployment.

The gap the research identifies is, at bottom, a documentation problem. When a model update causes harm or shifts behavior in a way that affects how millions of people are served, screened, or advised, there is no requirement that anyone produce a provable record of what the system did before the change and what it did after. Without that baseline, accountability is retrospective at best and absent at worst. Proportional oversight starts with knowing what changed.

Reported impact

Affected parties
Not publicly disclosed
Harm type
Not publicly disclosed
Scale
Not publicly disclosed
Financial impact
Not publicly disclosed
Regulatory action
Not publicly disclosed

Classification

Organization
Not publicly disclosed
AI system
Not publicly disclosed
Industry
Not publicly disclosed
Country
Not publicly disclosed
Provider
Not publicly disclosed
Incident type
Not publicly disclosed

Relevant governance controls

Governance control mapping is not available for this record.

  • No controls mappedNot publicly disclosed

Control mapping is analytical. It does not state that any control would have prevented the incident.

Sources and evidence

This record was researched and written by the Index. The event is also catalogued in the following database, which is listed for cross-reference.

OECD Wonk
Also catalogued in
Hundreds of AI Model Updates Go Live Every Year With Almost No Safety Review
2025-05-29T07:05:03