A 20-Million-Query Audit Found ChatGPT Treats Rich Countries Better Than Poor Ones
What happened
A large-scale audit of 20 million ChatGPT queries, published in February 2026, found a consistent pattern across geography, language, and topic: the model assigns positive traits like "intelligence" and "safety" to wealthy Western nations while placing low-income countries at the bottom of its rankings. The finding was not occasional or context-specific. Researchers described it as systematic amplification of existing global inequalities.
Researchers named the phenomenon "silicon gaze." The model's outputs, they argued, reflect a worldview inherited from its training data, shaped by the priorities of developers, platform owners, and the bodies of text used to build the model. ChatGPT learns predominantly from Western, high-income, and English-language sources, and it mirrors those imbalances in its answers. Ask it which countries are innovative or where investment flows best, and the answers correlate closely with historical wealth and geopolitical power.
The concern is not that a language model holds preferences. It is that those preferences are invisible to most users. When a student, analyst, or business leader queries the system for a neutral summary of global conditions, they receive no disclosure that the answer systematically favors certain countries. The model presents rankings as if they derive from evidence rather than an uneven historical record, which the researchers described as a process for "laundering" old prejudices through a high-technology interface.
The consequences compound. If a system used for investment screening, academic research, or infrastructure planning consistently marks wealthy nations as safer or more capable, those assessments feed back into decisions that concentrate resources where resources already sit. Countries underrepresented in digital archives, research publications, and English-language media start each query at a disadvantage the model amplifies rather than corrects. The audit placed this finding on a scale difficult to dismiss: 20 million queries, a consistent direction.
The accountability gap sits between the model's outputs and any verifiable standard for how it arrived at them. A user receiving a country ranking or cultural characterization has no way to inspect the training distribution that shaped the answer, no audit trail showing which sources received the most weight, and no provable record of what a system did when asked questions with significant geopolitical stakes. Without that record, structural bias reproduces itself at scale with no mechanism to catch or correct it.
Reported impact
- Affected parties
- Not publicly disclosed
- Harm type
- Not publicly disclosed
- Scale
- Not publicly disclosed
- Financial impact
- Not publicly disclosed
- Regulatory action
- Not publicly disclosed
Classification
Relevant governance controls
Governance control mapping is not available for this record.
- No controls mapped
Not publicly disclosed
Control mapping is analytical. It does not state that any control would have prevented the incident.
Sources and evidence
This record was researched and written by the Index. The event is also catalogued in the following database, which is listed for cross-reference.