Submit incident
Documented

Adobe Firefly Fixed Racial Bias by Breaking Historical Accuracy

January 1, 2024
Curated by Team Raidu · Reviewed by Shiva Ganesh
aiaaic:AIAAIC1391View source ↗
LinkedInX

What happened

In March 2024, tests by Semafor, the Daily Mail, and the New York Post found Adobe's Firefly image generator producing images that rewrote recorded history. The tool was generating Black and Asian soldiers depicted under a German flag in Nazi-era contexts, Black figures placed into scenes of the 1787 Constitutional Convention alongside the US Founding Fathers, Black Vikings, and a female Black Pope. The outputs came from straightforward historical queries, not adversarial prompts designed to break the system.

The pattern held across multiple outlets and multiple prompt types, which ruled out edge-case failures. Whatever diversification rule Firefly was applying, it was applying it uniformly: historical events, historical figures, and historical settings all received the same treatment regardless of what the documented record showed.

Adobe's designers were most likely trying to solve a real problem. Image generators trained on large datasets tend to default to all-white outputs when generating soldiers, politicians, or religious figures, because the training data reflects those demographic skews. The intended fix, applying racial diversification to generated images, has no obvious failure mode when applied to contemporary or fictional contexts. Applied to documented historical events, it produces the opposite of accuracy. The system had no mechanism to distinguish between a context where diversity correction is appropriate and one where the historical record is specific and documented.

Google had shut down Gemini's image tool weeks earlier for the same class of error, drawing similar criticism. Adobe noted that Firefly used a different dataset, trained on licensed stock images rather than the open web. The company's official position was that Firefly is not meant for photorealistic depictions of real or historical events, which was technically accurate but did not address why the tool was producing them by default when asked.

The incident points to a gap in how image generation systems are documented before deployment. There is no standard requirement to record which constraints a model applies to a given output type, how those constraints were tested, or who approved them. A user receiving Firefly's output had no means to inspect what rule produced the racial diversification, whether it was intentional, or whether it could be adjusted. A provable record of what a system did, which parameters shaped it, and who signed off on those parameters, would have made the design decision visible before a newspaper ran a test. Without that record, the first public evidence of the decision was the output itself.

Reported impact

Affected parties
Not publicly disclosed
Harm type
Not publicly disclosed
Scale
Not publicly disclosed
Financial impact
Not publicly disclosed
Regulatory action
Not publicly disclosed

Classification

Organization
Not publicly disclosed
AI system
Not publicly disclosed
Industry
Not publicly disclosed
Country
Not publicly disclosed
Provider
Not publicly disclosed
Incident type
Not publicly disclosed

Relevant governance controls

Governance control mapping is not available for this record.

  • No controls mappedNot publicly disclosed

Control mapping is analytical. It does not state that any control would have prevented the incident.

Sources and evidence

This record was researched and written by the Index. The event is also catalogued in the following database, which is listed for cross-reference.

AIAAIC Repository
Also catalogued in
Adobe Firefly Fixed Racial Bias by Breaking Historical Accuracy
2024