Submit incident
Documented

Decades-Old Student Data Was Put Up for Sale to AI Companies Without the Students' Knowledge

January 1, 2024
Curated by Team Raidu · Reviewed by Shiva Ganesh
aiaaic:AIAAIC1403View source ↗
LinkedInX

What happened

In February 2024, a company called Catalyst Research Alliance was found soliciting buyers for a university-linked academic dataset, pricing licenses at $25,000 each. The package it was selling contained 65 speech events, 85 hours of audio recordings, and 829 student papers, all drawn from collections associated with the University of Michigan. Catalyst described itself as a partner to both UM and North Carolina State University. Neither institution had authorized the sale, and the students whose work and voices were being sold had no idea it was happening.

The data came from two well-known academic corpora: MICASE, the Michigan Corpus of Academic Spoken English, and MICUSP, the Michigan Corpus of Upper-Level Student Papers. These collections were assembled over roughly a decade starting in 1997, with the explicit goal of supporting research into writing and articulation in academic settings. They were made freely available to other researchers in education for that purpose. That openness made them easy to locate, package, and pitch to organizations building language models, with no gatekeeping required to access what Catalyst was now trying to charge $25,000 to license.

The University of Michigan responded publicly once media coverage surfaced the offers. The university confirmed that students had given signed consent when the original research studies ran, but drew a clear line: that consent applied to participation in academic research, not to commercial licensing to AI companies years later. UM also confirmed it had cut ties with Catalyst Research Alliance entirely, ending any claim the company had to act under the university's name or authority.

What the incident illustrates is a governance gap that extends well beyond this case. Academic corpora get released for research use and then persist in accessible form indefinitely. Once they leave institutional control, no inherent mechanism prevents a downstream actor from repackaging the material and selling it for a purpose the original participants never agreed to. The consent forms signed by students between 1997 and 2007 were not written for the AI training economy of 2024, and nothing in the chain of custody required Catalyst to check whether the use it was proposing fell within the original terms.

The structural problem is that data released under one authorization rarely carries a verifiable record of what that authorization actually covered. A provable record of what a system was permitted to do with specific data, tied to the original consent terms and visible to anyone claiming to hold a license, would have made Catalyst's pitch immediately auditable rather than discoverable only through journalism. Without that kind of accountability trail, the gap between what participants consented to and what ultimately happens to their data stays invisible until something goes wrong.

Reported impact

Affected parties
Not publicly disclosed
Harm type
Not publicly disclosed
Scale
Not publicly disclosed
Financial impact
Not publicly disclosed
Regulatory action
Not publicly disclosed

Classification

Organization
Not publicly disclosed
AI system
Not publicly disclosed
Industry
Not publicly disclosed
Country
Not publicly disclosed
Provider
Not publicly disclosed
Incident type
Not publicly disclosed

Relevant governance controls

Governance control mapping is not available for this record.

  • No controls mappedNot publicly disclosed

Control mapping is analytical. It does not state that any control would have prevented the incident.

Sources and evidence

This record was researched and written by the Index. The event is also catalogued in the following database, which is listed for cross-reference.

AIAAIC Repository
Also catalogued in
Decades-Old Student Data Was Put Up for Sale to AI Companies Without the Students' Knowledge
2024