NewsTradingSentimentEventsCommunityBriefing
Tech

AI Chatbots Fail Financial Queries 57% of the Time, Study Finds

By Tech Desk · · 2 min read
A server rack with blinking status lights
Illustration: Tradingbird, based on a photo published by InvestmentNews

A new benchmark shows major AI models give wrong or incomplete money advice more often than right, with free tools performing worst.

Key points

  • Average AI accuracy for financial questions is 43%, with models failing 57% of the time.
  • Accuracy drops to 12% on complex multi-part financial scenarios involving tax rules.
  • Free AI models fail 63% of the time, significantly worse than paid models at 49%.

Popular AI chatbots provide incorrect or incomplete answers to financial questions more often than they get them right, according to a new comprehensive study. The research, conducted by Saturn, a UK-based technology firm serving thousands of financial advisors, reveals that the average accuracy rate across leading models is just 43%.

This means that for more than half of all queries, users receive guidance that is flawed or missing critical details. The study assessed over 10,000 responses from 18 different AI models, testing them against 121 specific money-related scenarios. The findings highlight a significant gap between consumer expectations and the actual reliability of these automated tools.

Complex questions expose major gaps

Performance drops sharply when queries involve complex calculations or interacting tax rules. On difficult multi-part scenarios, the average accuracy fell to just 12%, meaning the models made mistakes in nearly nine out of ten cases. Even on basic questions that require no calculation, accuracy averaged only 54%, indicating that fundamental financial literacy is not guaranteed.

The most common type of error was providing an incomplete answer, which accounted for over a third of all failures. Other significant issues included leaving out required figures or deadlines, fabricating rules that do not exist, and citing outdated regulatory guidance. These errors are particularly dangerous because they are often presented in a fluent, authoritative tone that mimics correct advice.

Free tools are less reliable

There is a clear trade-off between cost and accuracy in the AI market. Free versions of chatbots failed 63% of the time, compared to 49% for paid-for models. The worst-performing free model produced wrong or incomplete answers 82% of the time, while the best free option still returned incorrect responses more than half the time.

Even the top-performing paid model, Claude Opus 5, achieved a pass rate of only 61%. This means it still failed to answer correctly in almost four out of every ten cases. On hard questions, even this top performer made mistakes 67% of the time, suggesting that no current AI model is fully reliable for high-stakes financial decisions.

Risks for low-income users

The divergence in performance creates a troubling equity issue. Consumers who cannot afford professional financial advice are often the same group most likely to rely on free AI tools. As a result, these users are exposed to the least accurate guidance, which risks compounding financial inequality rather than narrowing it.

InvestmentNews notes that this study underscores the limitations of using AI as a substitute for professional advice. With nearly half of consumers already using AI for saving and investing decisions, the high error rate suggests that users must remain cautious and verify critical information independently.

Based on reporting by InvestmentNews, compiled by the Tradingbird desk.

Read next

More in Tech

More from the Tech desk

All desk stories