Submit incident
Documented

Nvidia Trained NeMo on 196,000 Pirated Books and Pulled the Platform Without Explanation

January 1, 2024
Curated by Team Raidu · Reviewed by Shiva Ganesh
aiaaic:AIAAIC1427View source ↗
LinkedInX

What happened

In October 2023, Nvidia quietly withdrew its NeMo generative AI platform from public access and said almost nothing about why. By March 2024, three authors had filed a class action lawsuit that made the reason specific: NeMo had been trained on Books3, a dataset assembled from approximately 196,640 pirated books, and none of the writers whose work ended up in it had been asked, licensed, or compensated.

Brian Keene, Abdi Nazemian, and Stewart O'Nan filed the complaint alleging copyright infringement, arguing that their work had been copied into Books3 and used to train Nvidia's NeMo models without permission. Their filing described how Books3 was built by reproducing all of Bibliotek, a shadow library that had circulated as part of The Pile, a larger open-source training dataset previously hosted on AI community platform Hugging Face. The Pile had already been removed from Hugging Face following an earlier copyright complaint, but its constituent datasets, Books3 among them, had been downloaded and redistributed widely before the takedown.

The authors asked the court for financial compensation for the unauthorized use of their creative work and demanded the destruction of every copy of the Books3 dataset. They also pointed directly at Nvidia's pre-lawsuit behavior as evidence. When the NeMo platform came down in October 2023, Nvidia acknowledged in a statement that the model had been trained on a dataset containing "approximately" 196,640 books, a number that matched Books3 precisely. The plaintiffs treated that phrasing as an implicit concession of exactly the connection they were alleging.

Nvidia positioned its use of the material as a fair use question, a doctrine in US law that permits limited use of copyrighted material without a license. The case sat alongside a cluster of similar suits filed against other AI developers during the same period, including actions brought by other authors against OpenAI, reflecting the broader unresolved collision between large-scale model training and copyright law.

The core problem the incident surfaces is not unique to Nvidia. When a model is trained on a composite dataset assembled from multiple layers of sources and sub-datasets, the chain of provenance is rarely recorded in any form that makes accountability after the fact possible. A provable record of what a system was trained on, and whether each component carried cleared rights, would have made the dispute resolvable before deployment rather than through litigation filed years after the training data was ingested.

Reported impact

Affected parties
Not publicly disclosed
Harm type
Not publicly disclosed
Scale
Not publicly disclosed
Financial impact
Not publicly disclosed
Regulatory action
Not publicly disclosed

Classification

Organization
Not publicly disclosed
AI system
Not publicly disclosed
Industry
Not publicly disclosed
Country
Not publicly disclosed
Provider
Not publicly disclosed
Incident type
Not publicly disclosed

Relevant governance controls

Governance control mapping is not available for this record.

  • No controls mappedNot publicly disclosed

Control mapping is analytical. It does not state that any control would have prevented the incident.

Sources and evidence

This record was researched and written by the Index. The event is also catalogued in the following database, which is listed for cross-reference.

AIAAIC Repository
Also catalogued in
Nvidia Trained NeMo on 196,000 Pirated Books and Pulled the Platform Without Explanation
2024