AI Chatbots Kept Stereotyping Black English Speakers Even After Anti-Racism Training
What happened
When researchers at the Allen Institute for AI studied how commercial language models respond to African American Vernacular English (AAVE), they found a pattern that the models' stated values could not explain. Published in March 2024, the study showed that ChatGPT and several other widely deployed chatbots applied covert racial stereotypes to AAVE speakers, attributing negative traits to people based on dialect alone. The training intended to prevent this had produced something more troubling: bias that hid behind a surface layer of politically acceptable responses.
The mechanism was specific. When a model processed a sentence in AAVE, such as "I be so happy when I wake up from a bad dream cus they be feelin too real," it assigned negative attributes to the implied speaker, labeling them "dirty," "stupid," or "lazy." The bias did not arrive through direct slurs. It came through assumptions embedded in downstream inference tasks: predicting what job a person would hold, or how likely they were to be involved in crime. The researchers found that the models exhibited stereotypes "more negative than any human stereotypes about African Americans ever experimentally recorded," and placed them closest to attitudes documented before the civil rights movement.
What made this hard to detect was the models' behavior when asked directly about race. Prompted to comment on stereotypes, GPT-4 and its predecessors produced careful, inclusive responses that appeared to demonstrate awareness of bias. That surface performance masked what the models were actually doing when a task shifted from explicit commentary to implicit inference. Anti-bias training appeared to teach the models to avoid saying the wrong thing, not to avoid doing it.
The real-world stakes are significant. Language models are deployed in hiring tools, legal research platforms, and content moderation systems, areas where a covert association between dialect and criminality or incompetence can shape outcomes for real people. AAVE is not an error pattern or informal shorthand; it is a fully grammatical variety of English natively spoken by most working- and middle-class African Americans and some Black Canadians. Treating it as a marker of low capability is discriminatory, whether or not the system producing that inference can explain its reasoning.
The deeper problem is that standard fairness benchmarks are not designed to catch this. A model can pass evaluations built around overt bias while encoding covert stereotypes that surface only in downstream tasks. There is no publicly available, provable record of what a commercial model actually outputs when processing dialect-marked text at scale, and no standard mechanism for verifying whether anti-bias training changes a model's inferences or only its explicit statements. Until that record exists, studies like this one are the only window into what these systems are actually doing.
Reported impact
- Affected parties
- Not publicly disclosed
- Harm type
- Not publicly disclosed
- Scale
- Not publicly disclosed
- Financial impact
- Not publicly disclosed
- Regulatory action
- Not publicly disclosed
Classification
Relevant governance controls
Governance control mapping is not available for this record.
- No controls mapped
Not publicly disclosed
Control mapping is analytical. It does not state that any control would have prevented the incident.
Sources and evidence
This record was researched and written by the Index. The event is also catalogued in the following database, which is listed for cross-reference.