California Let AI Write the Bar Exam, Then Had the Same Company Check the Work
What happened
The State Bar of California disclosed in April 2025 that it had used generative AI to develop multiple-choice questions for its February 2025 bar exam. The bar exam is the licensing threshold for every lawyer who practices in California, which makes its questions among the most consequential standardized content in American professional life. The disclosure came after examinees had already sat the test and reported problems with specific questions, some of which had content errors or ambiguous framing that made them difficult to answer correctly.
The process carried a structural conflict of interest from the start. Non-lawyers at the testing contractor used AI tools to draft the questions; the same company then validated what it had produced. No independent legal review separated the drafting phase from the quality-check phase. Critics, including law professors and bar-preparation experts, argued that this arrangement meant errors had to be caught by the same party that introduced them, rather than by a genuinely separate evaluator with subject-matter credentials.
The practical results were predictable. Examinees reported technical and content-related problems during the February sitting, and the State Bar's Competence-Based Examination committee recommended scoring adjustments afterward, an implicit acknowledgment that the questions could not be graded on their original terms. For the people who sat that exam, a scoring adjustment resolves nothing cleanly. It does not tell them whether their result reflects their legal knowledge or the quality of the questions they received.
The bar's stated rationale was that California's Supreme Court had encouraged it to consider new technologies in the examination process. That guidance did not come with a requirement to skip validation steps, and deploying AI-assisted question development without independent legal review is not what responsible exploration of that guidance was meant to authorize. The gap between piloting a technology and using it in a high-stakes credentialing context without independent safeguards is consequential, not a procedural footnote.
What the incident exposes is a documentation and verification gap that no scoring adjustment closes. There is no public record of which questions were AI-generated, what prompts produced them, who reviewed them with what qualifications, or whether any independent legal expert checked the final question bank before the exam ran. A provable record of what a system produced, who evaluated it, and with what authority would have surfaced the conflict-of-interest structure before candidates sat the exam rather than after the bar had to quietly revise its scoring.
Reported impact
- Affected parties
- Not publicly disclosed
- Harm type
- Not publicly disclosed
- Scale
- Not publicly disclosed
- Financial impact
- Not publicly disclosed
- Regulatory action
- Not publicly disclosed
Classification
Relevant governance controls
Governance control mapping is not available for this record.
- No controls mapped
Not publicly disclosed
Control mapping is analytical. It does not state that any control would have prevented the incident.
Sources and evidence
This record was researched and written by the Index. The event is also catalogued in the following database, which is listed for cross-reference.