Decades-Old Student Data Was Put Up for Sale to AI Companies Without the Students' Knowledge
What happened
In February 2024, a company called Catalyst Research Alliance was found soliciting buyers for a university-linked academic dataset, pricing licenses at $25,000 each. The package it was selling contained 65 speech events, 85 hours of audio recordings, and 829 student papers, all drawn from collections associated with the University of Michigan. Catalyst described itself as a partner to both UM and North Carolina State University. Neither institution had authorized the sale, and the students whose work and voices were being sold had no idea it was happening.
The data came from two well-known academic corpora: MICASE, the Michigan Corpus of Academic Spoken English, and MICUSP, the Michigan Corpus of Upper-Level Student Papers. These collections were assembled over roughly a decade starting in 1997, with the explicit goal of supporting research into writing and articulation in academic settings. They were made freely available to other researchers in education for that purpose. That openness made them easy to locate, package, and pitch to organizations building language models, with no gatekeeping required to access what Catalyst was now trying to charge $25,000 to license.
The University of Michigan responded publicly once media coverage surfaced the offers. The university confirmed that students had given signed consent when the original research studies ran, but drew a clear line: that consent applied to participation in academic research, not to commercial licensing to AI companies years later. UM also confirmed it had cut ties with Catalyst Research Alliance entirely, ending any claim the company had to act under the university's name or authority.
What the incident illustrates is a governance gap that extends well beyond this case. Academic corpora get released for research use and then persist in accessible form indefinitely. Once they leave institutional control, no inherent mechanism prevents a downstream actor from repackaging the material and selling it for a purpose the original participants never agreed to. The consent forms signed by students between 1997 and 2007 were not written for the AI training economy of 2024, and nothing in the chain of custody required Catalyst to check whether the use it was proposing fell within the original terms.
The structural problem is that data released under one authorization rarely carries a verifiable record of what that authorization actually covered. A provable record of what a system was permitted to do with specific data, tied to the original consent terms and visible to anyone claiming to hold a license, would have made Catalyst's pitch immediately auditable rather than discoverable only through journalism. Without that kind of accountability trail, the gap between what participants consented to and what ultimately happens to their data stays invisible until something goes wrong.
Reported impact
- Affected parties
- Not publicly disclosed
- Harm type
- Not publicly disclosed
- Scale
- Not publicly disclosed
- Financial impact
- Not publicly disclosed
- Regulatory action
- Not publicly disclosed
Classification
Relevant governance controls
Governance control mapping is not available for this record.
- No controls mapped
Not publicly disclosed
Control mapping is analytical. It does not state that any control would have prevented the incident.
Sources and evidence
This record was researched and written by the Index. The event is also catalogued in the following database, which is listed for cross-reference.