The Dataset Built to Block Nudity Was Hiding Illegal Images of Children
What happened
NudeNet was presented as infrastructure for safety. It contained more than 700,000 images scraped from the internet, assembled specifically so that developers could train AI classifiers to detect and filter nudity automatically. In December 2025, an analysis by the Canadian Centre for Child Protection found that the dataset had been hosting nearly 680 images of confirmed or suspected child sexual abuse material since at least June 2019.
The images included photographs of known victims, depictions of the genital and anal areas of pre-pubescent and post-pubescent children, and images of sexual or abusive acts involving children and teenagers. The dataset was publicly available on Academic Torrents for more than six years before the discovery. During that time, it was cited in over 250 published academic works by researchers building the exact kind of AI classifiers that content platforms rely on to catch illegal material. Following the Canadian Centre's findings, a removal notice was issued and the images were taken down from the hosting service.
The contamination happened because the collection was assembled by scraping social media and pornographic websites at scale with no meaningful ethical review and no systematic verification of what the images actually contained. Speed and volume were the implicit priorities, and the assumption was that the source material was legal. NudeNet was not unusual in this regard. LAION-5B, one of the most widely used image training datasets in the world, was later found to contain more than 1,000 verified instances of the same category of illegal material, the result of the same scrape-first logic applied at even greater scale.
The downstream implications compound the original harm. Victims whose images entered NudeNet were re-exposed through every training run and every model the dataset touched. AI systems trained on the contaminated data could inherit patterns derived from illegal imagery, undermining the stated purpose of building safer content detection. Researchers who cited the dataset in good faith now face potential legal liability for work built on material that was never legal to possess or distribute.
Nothing in the record suggests there was any system in place to verify the provenance of images before they entered the collection, or to flag anomalous content after the fact. A provable record of what a system ingested, where each image came from, and who reviewed the collection before it was made public would not have prevented the scraping, but it would have surfaced the contamination years earlier and created an accountability trail for every downstream use. The dataset circulated for six years without one.
Reported impact
- Affected parties
- Not publicly disclosed
- Harm type
- Not publicly disclosed
- Scale
- Not publicly disclosed
- Financial impact
- Not publicly disclosed
- Regulatory action
- Not publicly disclosed
Classification
Relevant governance controls
Governance control mapping is not available for this record.
- No controls mapped
Not publicly disclosed
Control mapping is analytical. It does not state that any control would have prevented the incident.
Sources and evidence
This record was researched and written by the Index. The event is also catalogued in the following database, which is listed for cross-reference.