Submit incident
Documented

Reddit Drew a Line Around Its Data, but Had No Way to Audit Who Had Already Crossed It

January 1, 2024
Curated by Team Raidu · Reviewed by Shiva Ganesh
aiaaic:AIAAIC1477View source ↗
LinkedInX

What happened

In April 2024, Reddit COO Jen Wong stated publicly that AI companies should not use data from the platform for commercial purposes without permission. The statement was not a cease-and-desist letter. It was a warning, and the distinction matters: Reddit was asserting a right it had not yet fully tested against the companies most likely to push back.

Wong's position was precise. AI firms scraping Reddit's content without authorization risked committing copyright infringement, in her framing, and potentially manipulating Reddit's systems in ways that violated the platform's terms of service. The statement named no company directly, but the targets were not ambiguous. Large AI labs had been drawing on Reddit's archive of human conversation for years as raw material for language model training, frequently without any licensing agreement in place. Reddit's posts and comment threads, spanning two decades and hundreds of millions of users, are among the densest concentrations of conversational text on the public internet.

Reddit's timing was not incidental. The company was preparing for a public offering, and the monetization of its data through paid API access and formal licensing deals with select AI firms was central to its pitch to investors. A platform that could not control who accessed its content was one with a weakened IP position going into markets. Wong's warning served as both a legal boundary statement and a signal that Reddit intended to take enforcement seriously, at least going forward.

The difficulty is that the warning landed in a landscape with no reliable mechanism for tracing what had already happened. AI training pipelines ingest data at scale from hundreds of sources simultaneously, and the companies running those pipelines face no requirement to disclose what they used, when, or in what volume. A lab could have gathered years of Reddit threads before April 2024 and Reddit would have had no ready means to prove the scope of it, let alone seek meaningful remediation. The warning addressed future conduct; it had nothing binding to say about the past.

That gap is what this record actually documents. Without a provable record of what a system ingested, from where, and under what terms, a platform's assertion of data rights becomes a negotiating posture rather than an enforceable claim. Reddit could warn, but it could not verify. The infrastructure that would make these disputes resolvable on facts rather than inference, a traceable, auditable account of what entered a model's training data and when, does not exist at the scale these companies operate, which means every future warning faces the same problem this one did.

Reported impact

Affected parties
Not publicly disclosed
Harm type
Not publicly disclosed
Scale
Not publicly disclosed
Financial impact
Not publicly disclosed
Regulatory action
Not publicly disclosed

Classification

Organization
Not publicly disclosed
AI system
Not publicly disclosed
Industry
Not publicly disclosed
Country
Not publicly disclosed
Provider
Not publicly disclosed
Incident type
Not publicly disclosed

Relevant governance controls

Governance control mapping is not available for this record.

  • No controls mappedNot publicly disclosed

Control mapping is analytical. It does not state that any control would have prevented the incident.

Sources and evidence

This record was researched and written by the Index. The event is also catalogued in the following database, which is listed for cross-reference.

AIAAIC Repository
Also catalogued in
Reddit Drew a Line Around Its Data, but Had No Way to Audit Who Had Already Crossed It
2024