Apple Trained Its AI on Pirated Books Without Telling the Authors
What happened
When Apple announced Apple Intelligence and its market value jumped by more than two hundred billion dollars in a single day, none of that gain reached the authors whose work had apparently made it possible. Two of those authors decided to do something about it.
Susana Martinez-Conde and Stephen Macknik, both professors at SUNY Downstate Health Sciences University, filed a proposed class action lawsuit in a California federal court in October 2025, alleging that Apple had used copyrighted material without permission to train its Apple Intelligence AI model. The two books at the center of the complaint are "Champions of Illusion: The Science Behind Mind-Boggling Images and Mystifying Brain Puzzles" and "Sleights of Mind: What the Neuroscience of Magic Reveals About Our Everyday Deceptions." Neither author consented to that use.
The mechanism the complaint targets is Books3, a dataset assembled from what are commonly called shadow libraries: collections of pirated books aggregated and circulated online. Apple allegedly sourced thousands of copyrighted titles from Books3 and similar repositories to expand its training data, bypassing the licensing negotiations, royalty arrangements, and consent requirements that would have applied in any conventional publishing deal. The complaint also alleges that Apple scraped additional copyrighted materials directly from the web.
Martinez-Conde v. Apple is one entry in a growing list of similar suits filed against major technology companies. Authors and publishers have brought parallel actions against OpenAI, Microsoft, and Meta on similar copyright infringement grounds. The broader pattern is consistent: AI developers building at speed needed data at scale, and shadow libraries offered a ready shortcut. The authors whose work filled those libraries were never consulted and have no visibility into how their writing was used, in what quantities, or to support which product features.
The case points to a structural gap in how AI training pipelines are documented. No external party can currently verify which books appeared in a company's training corpus, in what volume, or when. Without a provable record of what a system was trained on and what rights were cleared before that training ran, claims of infringement rest on circumstantial evidence rather than auditable fact. The authors seek monetary damages and an injunction to stop continued unauthorized use, but neither outcome resolves the underlying absence: an accountability trail that would have made this dispute either preventable or immediately verifiable when it arose.
Reported impact
- Affected parties
- Not publicly disclosed
- Harm type
- Not publicly disclosed
- Scale
- Not publicly disclosed
- Financial impact
- Not publicly disclosed
- Regulatory action
- Not publicly disclosed
Classification
Relevant governance controls
Governance control mapping is not available for this record.
- No controls mapped
Not publicly disclosed
Control mapping is analytical. It does not state that any control would have prevented the incident.
Sources and evidence
This record was researched and written by the Index. The event is also catalogued in the following database, which is listed for cross-reference.