Nvidia Trained NeMo on 196,000 Pirated Books and Pulled the Platform Without Explanation
What happened
In October 2023, Nvidia quietly withdrew its NeMo generative AI platform from public access and said almost nothing about why. By March 2024, three authors had filed a class action lawsuit that made the reason specific: NeMo had been trained on Books3, a dataset assembled from approximately 196,640 pirated books, and none of the writers whose work ended up in it had been asked, licensed, or compensated.
Brian Keene, Abdi Nazemian, and Stewart O'Nan filed the complaint alleging copyright infringement, arguing that their work had been copied into Books3 and used to train Nvidia's NeMo models without permission. Their filing described how Books3 was built by reproducing all of Bibliotek, a shadow library that had circulated as part of The Pile, a larger open-source training dataset previously hosted on AI community platform Hugging Face. The Pile had already been removed from Hugging Face following an earlier copyright complaint, but its constituent datasets, Books3 among them, had been downloaded and redistributed widely before the takedown.
The authors asked the court for financial compensation for the unauthorized use of their creative work and demanded the destruction of every copy of the Books3 dataset. They also pointed directly at Nvidia's pre-lawsuit behavior as evidence. When the NeMo platform came down in October 2023, Nvidia acknowledged in a statement that the model had been trained on a dataset containing "approximately" 196,640 books, a number that matched Books3 precisely. The plaintiffs treated that phrasing as an implicit concession of exactly the connection they were alleging.
Nvidia positioned its use of the material as a fair use question, a doctrine in US law that permits limited use of copyrighted material without a license. The case sat alongside a cluster of similar suits filed against other AI developers during the same period, including actions brought by other authors against OpenAI, reflecting the broader unresolved collision between large-scale model training and copyright law.
The core problem the incident surfaces is not unique to Nvidia. When a model is trained on a composite dataset assembled from multiple layers of sources and sub-datasets, the chain of provenance is rarely recorded in any form that makes accountability after the fact possible. A provable record of what a system was trained on, and whether each component carried cleared rights, would have made the dispute resolvable before deployment rather than through litigation filed years after the training data was ingested.
Reported impact
- Affected parties
- Not publicly disclosed
- Harm type
- Not publicly disclosed
- Scale
- Not publicly disclosed
- Financial impact
- Not publicly disclosed
- Regulatory action
- Not publicly disclosed
Classification
Relevant governance controls
Governance control mapping is not available for this record.
- No controls mapped
Not publicly disclosed
Control mapping is analytical. It does not state that any control would have prevented the incident.
Sources and evidence
This record was researched and written by the Index. The event is also catalogued in the following database, which is listed for cross-reference.