Tax Chatbots from TurboTax and H&R Block Were Wrong More Often Than Right
What happened
When Intuit and H&R Block added AI chatbots to their tax-preparation software in early 2024, the pitch was simple: get answers to tax questions without waiting for a human adviser. What Washington Post technology journalist Geoffrey A. Fowler found when he tested both systems alongside two tax experts was rather different. The chatbots were wrong, misleading, or unhelpful at a rate that would be unremarkable in a trivia game but is genuinely dangerous when the stakes are a federal tax filing and a potential audit.
Fowler ran 16 test questions through Intuit's Intuit Assist chatbot and received inaccurate responses to more than half of them. H&R Block's system compounded the problem with a particular kind of wrongness: it did not hedge or decline to answer, it confidently recommended an incorrect filing status and gave a flat-out inaccurate description of IRS guidance on cryptocurrency reporting. Confident errors are worse than uncertain ones in high-stakes domains, because a user has no signal to prompt a second check. Both failures occurred during a period when millions of taxpayers were making consequential decisions about how to file.
Intuit updated Intuit Assist after receiving Fowler's findings, which is a reasonable response to a structured critique. The revised version performed better, but still proved unhelpful on a quarter of the questions it was given. One in four is not a rounding error for a product positioned as a reliable adviser on legal documents submitted to the federal government. It means a meaningful fraction of every real-world session could end with a user acting on guidance that did not hold up to scrutiny.
Fowler's published warning was direct: users should be "especially wary of generative AI when there are real-life consequences to it being wrong," and the specific consequence he named was a tax audit. His broader observation was sharper still. "We can't necessarily trust companies experimenting with AI to make the right decisions to protect our interests." Both systems were commercial products available to paying customers at the time of testing, not internal prototypes, yet neither had been held to any published external accuracy standard before launch.
There is no public record of what accuracy threshold, if any, either company required these systems to meet before embedding them in software used to file legal returns. A user who received a wrong answer had no way to know it was wrong at the time, no session log to produce if questioned, and no formal recourse tied to the specific failure. A provable record of what a system said, in what session, and how often it erred would make that exposure visible before the audit notice arrives rather than after.
Reported impact
- Affected parties
- Not publicly disclosed
- Harm type
- Not publicly disclosed
- Scale
- Not publicly disclosed
- Financial impact
- Not publicly disclosed
- Regulatory action
- Not publicly disclosed
Classification
Relevant governance controls
Governance control mapping is not available for this record.
- No controls mapped
Not publicly disclosed
Control mapping is analytical. It does not state that any control would have prevented the incident.
Sources and evidence
This record was researched and written by the Index. The event is also catalogued in the following database, which is listed for cross-reference.