OpenAI Built a Marketplace for Custom Chatbots Without Building a Way to Keep Stolen Content Out
What happened
OpenAI's GPT Store launched with a clear premise: developers could upload their own data, configure a custom chatbot for a specific use case, and distribute it to anyone. The upload step was frictionless by design. What the platform did not build, before opening to the public in April 2024, was a reliable way to check whether the data being uploaded belonged to the person uploading it.
The gap became visible when a Danish textbook publisher went public with what it had found. Blichfeldt Andersen, Publishing Director at the company, told WIRED that third-party developers were regularly pulling copyrighted educational materials into their custom bots by uploading them as training data. Andersen had identified specific violations and reported them to OpenAI directly. The company's response was not a policy change or a technical fix. It was a complaints queue.
That queue is where the problem concentrates. Andersen described the process for identifying and removing infringing bots as overly burdensome, a burden that falls entirely on the copyright holder rather than on the platform that accepted the upload in the first place. His company was effectively doing moderation work for a marketplace it had no role in building and no financial stake in running. He said that without meaningful improvements, the publisher was considering legal action.
The risk is not confined to one sector or one country. The GPT Store's developer base spans a wide range of technical and legal sophistication. Some builders are professional developers with legal teams. Many are not. The platform's default assumption, that anyone configuring a custom bot would understand and observe copyright restrictions, overstates what most people know about intellectual property when they are trying to make a tool that works. Nothing in the store's original design required a developer to attest to holding rights before a bot was published and made available to users.
What this episode exposes is not primarily a question of what any individual developer chose to upload. It is a question of what a platform chose not to verify before distributing the result. A provable record of what data went into a custom model, when it was uploaded, and whether the person uploading it confirmed they held the rights, would make violations easier to find, easier to attribute, and faster to act on. Without that record, the only people positioned to catch the problem are the ones who have already been harmed by it.
Reported impact
- Affected parties
- Not publicly disclosed
- Harm type
- Not publicly disclosed
- Scale
- Not publicly disclosed
- Financial impact
- Not publicly disclosed
- Regulatory action
- Not publicly disclosed
Classification
Relevant governance controls
Governance control mapping is not available for this record.
- No controls mapped
Not publicly disclosed
Control mapping is analytical. It does not state that any control would have prevented the incident.
Sources and evidence
This record was researched and written by the Index. The event is also catalogued in the following database, which is listed for cross-reference.