Hundreds of Contractors Were Reading ChatGPT Conversations, and Users Were Never Told
What happened
ChatGPT presents itself as a private dialogue between a user and a machine. The reality, documented in reporting and subsequent legal filings, is different. OpenAI employs hundreds of contractors whose job is to read real user conversations, flagging and rating exchanges to improve the model's behavior. The users on the other end of those conversations had no meaningful way to know that was happening.
The contractors are part of a standard practice in large-scale language model development called human feedback. Conversations are sampled, queued, and reviewed, often for tone, accuracy, and safety. Privacy filters are supposed to remove identifying information before a human ever sees a transcript, but those filters do not catch everything. Sensitive personal data, including medical information, relationship disclosures, and other details users typed believing they were interacting with a system, reached contractor screens.
The practice was not hidden in the sense of being technically secret. Terms of service disclosed data use in broad language. But disclosure buried in a legal document that few users read is not the same as informed understanding. Most people interacting with ChatGPT had no working model of what happened to their inputs after the conversation ended. The gap between what users believed and what was actually occurring is where the harm sits.
The exposure produced legal consequences. Lawsuits alleging harm were filed, with the adequacy of data protection measures as the central question. Whether privacy filters met reasonable standards, and whether consent language was sufficient to cover contractor review of sensitive personal disclosures, are now matters in dispute. Regulatory scrutiny of how AI companies handle training data has accelerated in several jurisdictions, driven in part by incidents like this one.
The documentation gap at the center of this situation is not a technical problem. Logging which conversations were reviewed, by whom, under what data protection controls, and what happened to flagged material is entirely feasible. Without that log, users whose conversations were handled cannot know whether their data was processed correctly, and no external auditor can verify that stated privacy filters worked as described. A provable record of what a system did with user-submitted content, maintained and accessible for oversight, is the only thing that turns a privacy policy commitment into something that can actually be checked.
Reported impact
- Affected parties
- Not publicly disclosed
- Harm type
- Not publicly disclosed
- Scale
- Not publicly disclosed
- Financial impact
- Not publicly disclosed
- Regulatory action
- Not publicly disclosed
Classification
Relevant governance controls
Governance control mapping is not available for this record.
- No controls mapped
Not publicly disclosed
Control mapping is analytical. It does not state that any control would have prevented the incident.
Sources and evidence
This record was researched and written by the Index. The event is also catalogued in the following database, which is listed for cross-reference.