Ten Million Articles Without a License: The Times Takes AI Training to Court
What happened
Ten million articles is not an accident. When the New York Times filed suit against OpenAI and Microsoft in federal court in New York, the complaint described a systematic use of the publication's entire archive, piece by piece, to train large language models sold to the public. The suit alleged that no license was obtained, no permission was sought, and no payment was made. What the Times was describing was not a scraping incident but a business model.
The core allegation is concrete: over 10 million copyrighted news articles from the Times were ingested into the training datasets behind models like ChatGPT. Journalism costs money to produce. Investigations, correspondents abroad, fact-checkers, editors, lawyers who review sensitive material before publication: all of that overhead is embedded in the value of what the Times publishes. The lawsuit argues that when a model learns from that material without a license, it extracts the value of the work without bearing any of the cost of producing it.
Microsoft is named alongside OpenAI because it has invested heavily in the same technology and distributes the resulting products through its own platforms and partnerships. That connection matters beyond the question of corporate liability. It shows that the supply chain of a major commercial AI deployment runs through content that the people who created it never agreed to hand over and were never compensated for.
The case sits alongside a pattern of similar disputes between AI developers and copyright holders. Publishers, authors, and creators across multiple industries have made versions of the same argument: that the training pipeline for generative AI was built substantially on material it was not licensed to use, and that the systems sold on the strength of that material represent an ongoing commercial profit drawn from an original taking that no contract authorized.
What makes cases like this hard to resolve is also what makes them important. There is no standard mechanism for a content creator to verify what went into a model's training data, which means there is no standard way to establish a violation without a lawsuit and the discovery process that comes with it. A provable record of what a system was trained on, when, and under what authority, would change the evidentiary baseline entirely. Right now the burden falls on the party that was never in the room when the decision was made.
Reported impact
- Affected parties
- Not publicly disclosed
- Harm type
- Not publicly disclosed
- Scale
- Not publicly disclosed
- Financial impact
- Not publicly disclosed
- Regulatory action
- Not publicly disclosed
Classification
Relevant governance controls
Governance control mapping is not available for this record.
- No controls mapped
Not publicly disclosed
Control mapping is analytical. It does not state that any control would have prevented the incident.
Sources and evidence
This record was researched and written by the Index. The event is also catalogued in the following database, which is listed for cross-reference.