Submit incident
Documented

One in Six Peer Reviews at Major AI Conferences Was Likely Written by AI

January 1, 2024
Curated by Team Raidu · Reviewed by Shiva Ganesh
aiaaic:AIAAIC1434View source ↗
LinkedInX

What happened

In March 2024, researchers from Stanford University, NEC Labs America, and UC Santa Barbara published a statistical analysis of peer reviews submitted to four major AI conferences: ICLR 2024, NeurIPS 2023, CoRL 2023, and EMNLP 2023. Their finding was specific and sourced: between 6.5 percent and 16.9 percent of submitted review text had likely been substantially modified by large language models, meaning the change went well beyond spell-checking or minor rewrites. At the high end of their estimate, roughly one in six reviews was at least partly a machine's output.

The researchers identified the pattern through vocabulary. Large language models reach for a predictable set of adjectives when writing evaluations: "commendable," "innovative," and "comprehensive" appear at rates statistically higher than in human-authored text. That signal was consistent enough across all four conference datasets that the team could build a reliable estimate of LLM involvement in the review pool. The authors acknowledged a range rather than a single figure, but the floor of that range, 6.5 percent, was large enough to be consequential on its own.

Peer review is the mechanism that determines which research gets accepted, which claims earn credibility, and which directions a field moves in next. For AI conferences specifically, a discipline that builds the very tools being misused here, a reviewer delegating their evaluation to a model means the submitting author receives feedback shaped by pattern-matching on surface features rather than by expert judgment. The research may be technically original; the review of it may not be.

None of the conferences had a policy at the time that formally detected or prohibited LLM use in peer review. The researchers surfaced the pattern only retrospectively, by analyzing text statistically after reviews had already been submitted, acted on, and used to decide acceptance. There was no checkpoint in the submission pipeline that would have caught it in real time, and reviewers had no obligation to disclose what tools they used.

This is a record-keeping and accountability problem as much as an integrity one. What each reviewer actually did, whether they read the paper carefully, consulted prior work, or handed the writing to a model, left no trace in any system the conference could audit. A provable record of what a system did in the review pipeline, not just the final text submitted, would make that delegation visible when it happens rather than months later, when researchers are reduced to counting adjectives to find out.

Reported impact

Affected parties
Not publicly disclosed
Harm type
Not publicly disclosed
Scale
Not publicly disclosed
Financial impact
Not publicly disclosed
Regulatory action
Not publicly disclosed

Classification

Organization
Not publicly disclosed
AI system
Not publicly disclosed
Industry
Not publicly disclosed
Country
Not publicly disclosed
Provider
Not publicly disclosed
Incident type
Not publicly disclosed

Relevant governance controls

Governance control mapping is not available for this record.

  • No controls mappedNot publicly disclosed

Control mapping is analytical. It does not state that any control would have prevented the incident.

Sources and evidence

This record was researched and written by the Index. The event is also catalogued in the following database, which is listed for cross-reference.

AIAAIC Repository
Also catalogued in
One in Six Peer Reviews at Major AI Conferences Was Likely Written by AI
2024