Submit incident
Documented

The Dataset Built to Block Nudity Was Hiding Illegal Images of Children

January 1, 2025
Curated by Team Raidu · Reviewed by Shiva Ganesh
aiaaic:AIAAIC2159View source ↗
LinkedInX

What happened

NudeNet was presented as infrastructure for safety. It contained more than 700,000 images scraped from the internet, assembled specifically so that developers could train AI classifiers to detect and filter nudity automatically. In December 2025, an analysis by the Canadian Centre for Child Protection found that the dataset had been hosting nearly 680 images of confirmed or suspected child sexual abuse material since at least June 2019.

The images included photographs of known victims, depictions of the genital and anal areas of pre-pubescent and post-pubescent children, and images of sexual or abusive acts involving children and teenagers. The dataset was publicly available on Academic Torrents for more than six years before the discovery. During that time, it was cited in over 250 published academic works by researchers building the exact kind of AI classifiers that content platforms rely on to catch illegal material. Following the Canadian Centre's findings, a removal notice was issued and the images were taken down from the hosting service.

The contamination happened because the collection was assembled by scraping social media and pornographic websites at scale with no meaningful ethical review and no systematic verification of what the images actually contained. Speed and volume were the implicit priorities, and the assumption was that the source material was legal. NudeNet was not unusual in this regard. LAION-5B, one of the most widely used image training datasets in the world, was later found to contain more than 1,000 verified instances of the same category of illegal material, the result of the same scrape-first logic applied at even greater scale.

The downstream implications compound the original harm. Victims whose images entered NudeNet were re-exposed through every training run and every model the dataset touched. AI systems trained on the contaminated data could inherit patterns derived from illegal imagery, undermining the stated purpose of building safer content detection. Researchers who cited the dataset in good faith now face potential legal liability for work built on material that was never legal to possess or distribute.

Nothing in the record suggests there was any system in place to verify the provenance of images before they entered the collection, or to flag anomalous content after the fact. A provable record of what a system ingested, where each image came from, and who reviewed the collection before it was made public would not have prevented the scraping, but it would have surfaced the contamination years earlier and created an accountability trail for every downstream use. The dataset circulated for six years without one.

Reported impact

Affected parties
Not publicly disclosed
Harm type
Not publicly disclosed
Scale
Not publicly disclosed
Financial impact
Not publicly disclosed
Regulatory action
Not publicly disclosed

Classification

Organization
Not publicly disclosed
AI system
Not publicly disclosed
Industry
Not publicly disclosed
Country
Not publicly disclosed
Provider
Not publicly disclosed
Incident type
Not publicly disclosed

Relevant governance controls

Governance control mapping is not available for this record.

  • No controls mappedNot publicly disclosed

Control mapping is analytical. It does not state that any control would have prevented the incident.

Sources and evidence

This record was researched and written by the Index. The event is also catalogued in the following database, which is listed for cross-reference.

AIAAIC Repository
Also catalogued in
The Dataset Built to Block Nudity Was Hiding Illegal Images of Children
2025