ChatGPT Screened Out Black Women From Tech Roles at Rates That Would Fail a US Discrimination Audit
What happened
In March 2024, Bloomberg researchers ran a controlled experiment to measure whether a widely deployed AI text system would sort job candidates differently depending on their apparent race and gender. They fed the system fictitious names and resumes grouped into four racial categories, White, Hispanic, Black, and Asian, and two gender categories, then asked it to evaluate candidates for four different job openings. The results were consistent and measurable: the system did not treat the applications equally.
GPT-3.5, then the most broadly used version of the model, repeatedly routed candidates with female names toward roles with historically higher concentrations of women, including HR positions. For technical roles like software engineer, the system chose Black women candidates 36 percent less frequently. Bloomberg calculated that patterns of this magnitude would fail the benchmarks the US applies when assessing job discrimination against protected groups under employment law.
The finding matters because it was not generated by adversarial prompting or edge-case inputs. Bloomberg was replicating what a recruitment professional might do: feed a stack of resumes into a generative AI system and ask it to help rank applicants. The model cooperated with that framing. It just did so while applying differential treatment that a human hiring manager, facing audit, would be required to explain and justify.
The experiment landed at a moment when AI tools are being adopted into automated hiring workflows at scale. Recruitment software vendors and HR teams have begun deploying generative models as a filter in the early screening phase, where large volumes of applications are winnowed before any human reads them. If those filters reproduce the biases Bloomberg documented, the bias becomes invisible and structural rather than an occasional judgment call a manager might catch and correct.
The documentation gap here is not theoretical. When a human screener passes over a candidate, that decision exists in a process that can be reviewed, challenged, or reversed. When a model makes the same pass at scale across thousands of applications, nothing in the standard pipeline requires anyone to log what the model did or why. That is exactly the kind of absence a provable record of what a system did is meant to address: not to prevent a model from making decisions, but to ensure each decision can be examined, audited, and held to the same standards the law already demands of human ones.
Reported impact
- Affected parties
- Not publicly disclosed
- Harm type
- Not publicly disclosed
- Scale
- Not publicly disclosed
- Financial impact
- Not publicly disclosed
- Regulatory action
- Not publicly disclosed
Classification
Relevant governance controls
Governance control mapping is not available for this record.
- No controls mapped
Not publicly disclosed
Control mapping is analytical. It does not state that any control would have prevented the incident.
Sources and evidence
This record was researched and written by the Index. The event is also catalogued in the following database, which is listed for cross-reference.