Cohere Is Accused of Taking Publishers' Work to Train, Reprinting It on Demand, and Inventing Stories Under Their Names
What happened
The lawsuit filed against Cohere in February 2025 is not a standard training-data dispute. Most AI copyright cases turn on whether scraping copyrighted material to build a model constitutes infringement at the point of ingestion. This one carries that claim, and then goes further. The consortium of major news publishers behind the filing alleges that Cohere's system also reproduced portions of their articles, and in some cases entire copies, as direct responses to user queries, turning a training complaint into a live distribution complaint.
The plaintiffs include Condé Nast, The Atlantic, Forbes Media, Guardian News and Media, Insider, and the Los Angeles Times, among others. Their complaint describes a pattern it calls "massive" and "systematic": a user asking the system a question about a topic could receive substantial text lifted from the publisher's own coverage rather than a summary or citation. At least 4,000 copyrighted works were allegedly used without permission to build the models in the first place, and then those same works were reproduced in output without any license to do so.
The second claim in the filing is distinct and, in some ways, more damaging to publishers. Cohere's system allegedly generated content that none of the plaintiff outlets had ever published, then attributed it to them by name. A user who asked about a specific news outlet's coverage could receive a fabricated story presented as genuine reporting from that outlet. This is not a training-data argument. It is an allegation of real-time misrepresentation, where a system invents journalism and assigns it a credible byline.
Taken together, the claims describe an AI product that consumed copyrighted journalism to train, reproduced that journalism in responses, and manufactured new falsehoods that borrowed the credibility of established newsrooms. For publishers already operating under revenue pressure from digital advertising, a system that answers queries with their content without routing any traffic back, while simultaneously inventing stories under their mastheads, attacks both the economic and reputational foundation of what the original reporting was worth.
The case points to a verification gap that neither the industry nor regulators have closed. There is no requirement that an AI company document which copyrighted works entered its training pipeline, when, under what terms, or whether any system outputs at inference time reproduced those works or fabricated text in a source's name. Without that record, publishers cannot audit at scale what was taken, what was reproduced, or whether invented stories attributed to their outlets are still circulating. A provable record of what a system ingested and what it generated on a given query would not settle the legal dispute, but it would make the harm visible without requiring years of litigation to establish it.
Reported impact
- Affected parties
- Not publicly disclosed
- Harm type
- Not publicly disclosed
- Scale
- Not publicly disclosed
- Financial impact
- Not publicly disclosed
- Regulatory action
- Not publicly disclosed
Classification
Relevant governance controls
Governance control mapping is not available for this record.
- No controls mapped
Not publicly disclosed
Control mapping is analytical. It does not state that any control would have prevented the incident.
Sources and evidence
This record was researched and written by the Index. The event is also catalogued in the following database, which is listed for cross-reference.