Submit incident
Documented

Apple Trained Its AI on Pirated Books Without Telling the Authors

January 1, 2025
Curated by Team Raidu · Reviewed by Shiva Ganesh
aiaaic:AIAAIC2061View source ↗
LinkedInX

What happened

When Apple announced Apple Intelligence and its market value jumped by more than two hundred billion dollars in a single day, none of that gain reached the authors whose work had apparently made it possible. Two of those authors decided to do something about it.

Susana Martinez-Conde and Stephen Macknik, both professors at SUNY Downstate Health Sciences University, filed a proposed class action lawsuit in a California federal court in October 2025, alleging that Apple had used copyrighted material without permission to train its Apple Intelligence AI model. The two books at the center of the complaint are "Champions of Illusion: The Science Behind Mind-Boggling Images and Mystifying Brain Puzzles" and "Sleights of Mind: What the Neuroscience of Magic Reveals About Our Everyday Deceptions." Neither author consented to that use.

The mechanism the complaint targets is Books3, a dataset assembled from what are commonly called shadow libraries: collections of pirated books aggregated and circulated online. Apple allegedly sourced thousands of copyrighted titles from Books3 and similar repositories to expand its training data, bypassing the licensing negotiations, royalty arrangements, and consent requirements that would have applied in any conventional publishing deal. The complaint also alleges that Apple scraped additional copyrighted materials directly from the web.

Martinez-Conde v. Apple is one entry in a growing list of similar suits filed against major technology companies. Authors and publishers have brought parallel actions against OpenAI, Microsoft, and Meta on similar copyright infringement grounds. The broader pattern is consistent: AI developers building at speed needed data at scale, and shadow libraries offered a ready shortcut. The authors whose work filled those libraries were never consulted and have no visibility into how their writing was used, in what quantities, or to support which product features.

The case points to a structural gap in how AI training pipelines are documented. No external party can currently verify which books appeared in a company's training corpus, in what volume, or when. Without a provable record of what a system was trained on and what rights were cleared before that training ran, claims of infringement rest on circumstantial evidence rather than auditable fact. The authors seek monetary damages and an injunction to stop continued unauthorized use, but neither outcome resolves the underlying absence: an accountability trail that would have made this dispute either preventable or immediately verifiable when it arose.

Reported impact

Affected parties
Not publicly disclosed
Harm type
Not publicly disclosed
Scale
Not publicly disclosed
Financial impact
Not publicly disclosed
Regulatory action
Not publicly disclosed

Classification

Organization
Not publicly disclosed
AI system
Not publicly disclosed
Industry
Not publicly disclosed
Country
Not publicly disclosed
Provider
Not publicly disclosed
Incident type
Not publicly disclosed

Relevant governance controls

Governance control mapping is not available for this record.

  • No controls mappedNot publicly disclosed

Control mapping is analytical. It does not state that any control would have prevented the incident.

Sources and evidence

This record was researched and written by the Index. The event is also catalogued in the following database, which is listed for cross-reference.

AIAAIC Repository
Also catalogued in
Apple Trained Its AI on Pirated Books Without Telling the Authors
2025