Submit incident
Documented

Devin Could Not Do the Job Its Maker Said It Could

January 1, 2025
Curated by Team Raidu · Reviewed by Shiva Ganesh
aiaaic:AIAAIC1883View source ↗
LinkedInX

What happened

Researchers at Answer.AI spent a month running Devin through engineering tasks and published what they found: 14 of 20 tasks failed, 3 were inconclusive, and 3 succeeded. Cognition AI had launched Devin as the world's first AI software engineer, a fully autonomous tool capable of handling development work end to end. The study did not find a software engineer. It found a system that consistently produced unusable output.

The specific failures give the numbers weight. Asked to deploy multiple applications to Railway, Devin could not identify the deployment target as unsuitable and spent over a day pursuing nonviable approaches. Web scraping tasks sent it into loops, cycling through HTML parsing attempts without stopping or changing course. Security reviews generated false positives in large numbers alongside vulnerabilities that did not exist. Across all three failure modes, the system continued working confidently after the work had already gone wrong.

Two factors explain the gap. Cognition AI appears to have shipped Devin without testing at the scale its claims required. A tool positioned as a replacement for human engineers needs benchmarks that look like actual engineering work, not curated demonstrations. The company also raised money from Founders Fund and Khosla Ventures, creating investor expectations that appear to have pushed release ahead of readiness. Public communications treated the product as solved rather than in progress.

The broader implication of Answer.AI's study is not confined to a single company. The case for fully autonomous AI development tools has largely relied on controlled environments and scripted demos. A month of open-ended real tasks reveals what those formats conceal: cascading failures when conditions vary, poor recovery when tasks diverge from expectations, and an inability to recognize when to stop. The conclusion that AI tools perform better as assistants than as autonomous replacements is not a minor calibration note. It challenges the commercial premise that drove Devin's launch.

What independent evaluations like this one still cannot provide is a provable record of what a system did across the full range of tasks it claimed to support. Without that record, large capability claims travel for months before anyone runs the tests that check them. Cognition AI's investors, customers, and the engineers it was marketed to all relied on representations no public benchmark had validated. That gap is not specific to Devin. It runs through the autonomous AI category as a whole.

Reported impact

Affected parties
Not publicly disclosed
Harm type
Not publicly disclosed
Scale
Not publicly disclosed
Financial impact
Not publicly disclosed
Regulatory action
Not publicly disclosed

Classification

Organization
Not publicly disclosed
AI system
Not publicly disclosed
Industry
Not publicly disclosed
Country
Not publicly disclosed
Provider
Not publicly disclosed
Incident type
Not publicly disclosed

Relevant governance controls

Governance control mapping is not available for this record.

  • No controls mappedNot publicly disclosed

Control mapping is analytical. It does not state that any control would have prevented the incident.

Sources and evidence

This record was researched and written by the Index. The event is also catalogued in the following database, which is listed for cross-reference.

AIAAIC Repository
Also catalogued in
Devin Could Not Do the Job Its Maker Said It Could
2025