Two Newsrooms Say OpenAI Trained on Their Paywalled Archives Without Permission
What happened
The Seattle Times and Newsday filed a copyright lawsuit against OpenAI and Microsoft in September 2026, alleging that both companies scraped their articles, including content locked behind their paywalls, to build the training datasets that power ChatGPT, Microsoft Copilot, and Bing AI. The newspapers say the use of their copyrighted material was unauthorized and that neither company sought or received permission before incorporating that journalism into commercial AI products.
The detail that carries the most weight in the complaint is the paywall allegation. Paywalled content is not indexed by search engines or freely visible to the public; it requires a paid subscription to access. The newspapers are arguing that OpenAI and Microsoft did not merely sweep up text that was lying in the open. They got through a gate that existed specifically to protect that content commercially. If the allegation holds, the companies bypassed a deliberate restriction rather than taking what any internet user could already reach.
The harm the newspapers describe is economic and structural. Newsrooms that license content to research institutions, academic aggregators, and syndication partners charge for that access because their journalism carries commercial value. When a model can reproduce the substance or reasoning of a paywalled article in response to a user query, it competes directly with the product the newsroom sells. The Seattle Times and Newsday argue that this substitution effect damages their business regardless of whether the model reproduces any article word for word.
This lawsuit joins a wave of similar legal challenges from publishers, including the New York Times, which filed its own suit against OpenAI and Microsoft in late 2023. The pattern across these cases is consistent: publishers allege that both companies treated web-accessible text as freely usable training material and paywalled text as an obstacle to route around, and that neither category required an explicit license or payment.
The accountability gap the case exposes is fundamental. No public record exists of which sources fed which model. Without that, the newspapers cannot prove exactly how much of their work was used, and the defendants can dispute the scope. What the AI industry has not yet built is a provable record of what a system did: which sources it ingested, when, and under what authorization. Without that record, disputes over training data can only be resolved through litigation, discovery, and sealed settlements rather than a transparent audit. The lawsuit becomes a stand-in for verification infrastructure that should have existed before the first article was scraped.
Reported impact
- Affected parties
- Not publicly disclosed
- Harm type
- Not publicly disclosed
- Scale
- Not publicly disclosed
- Financial impact
- Not publicly disclosed
- Regulatory action
- Not publicly disclosed
Classification
Relevant governance controls
Governance control mapping is not available for this record.
- No controls mapped
Not publicly disclosed
Control mapping is analytical. It does not state that any control would have prevented the incident.
Sources and evidence
This record was researched and written by the Index. The event is also catalogued in the following database, which is listed for cross-reference.