Major Newsrooms Tell a Federal Judge That AI Training Is Not Fair Use
What happened
The New York Times, the Chicago Tribune, and a coalition of news organizations filed a motion in a Manhattan federal court in September 2026, asking a judge to reject the primary defense Microsoft and OpenAI had raised in a copyright lawsuit. The defendants argued that training an AI system on copyrighted news articles constitutes fair use under U.S. copyright law. The plaintiffs want that argument struck before the case reaches trial.
Fair use permits limited, unlicensed use of protected material, but the news organizations argue the doctrine does not stretch to cover what happened here. Training an AI system means ingesting large volumes of articles, absorbing their content and style at scale, and commercializing the result. The plaintiffs contend the aggregate taking, done without consent or payment, stripped the original work of its commercial value while the defendants profited from it. The argument is not that any single article was reproduced in full; it is that the scale and commercial application of the use defeat the fair use claim.
The lawsuit names the harm to journalism directly. AI systems trained on newsroom archives can produce summaries, analysis, and news-adjacent text that reduces readers' reasons to visit original sources or pay for subscriptions. That shrinks the revenue that funds the reporting in the first place. The harm is not located in one act of copying; it is in what the trained system can do afterward, and fair use arguments that focus narrowly on the act of ingestion tend to miss that downstream effect.
The motion also exposes how the AI training industry set its defaults. The standard approach was to collect as much text as possible from the public web, treat fair use as a working assumption, and sort out the legal questions later. For news organizations, that posture meant losing licensing revenue they could have negotiated if authorization had been required before training began, not after the models were already deployed and commercially valuable.
There is a narrower documentation problem underneath the legal one. Once a model is trained, establishing exactly which articles were ingested, in what volume, and in what form requires records the developers hold and currently have no legal obligation to disclose. A provable record of what a system did during its training, logged at the time and open to independent verification, would make the underlying factual dispute answerable without years of pretrial discovery. Without it, publishers and courts are left arguing about process from the outside looking in.
Reported impact
- Affected parties
- Not publicly disclosed
- Harm type
- Not publicly disclosed
- Scale
- Not publicly disclosed
- Financial impact
- Not publicly disclosed
- Regulatory action
- Not publicly disclosed
Classification
Relevant governance controls
Governance control mapping is not available for this record.
- No controls mapped
Not publicly disclosed
Control mapping is analytical. It does not state that any control would have prevented the incident.
Sources and evidence
This record was researched and written by the Index. The event is also catalogued in the following database, which is listed for cross-reference.