AI 'in most cases' gives incorrect answers to financial questions, study finds

Saturn Study: Chatbots Provide Incorrect Financial Answers “in Most Cases” / Photo: Stockinq / Shutterstock
A study by the technology company Saturn, as reported by the Financial Times, found that using AI-powered chatbots to handle financial matters can lead to significant losses. On average, in 57% of cases, the most popular AI models—ChatGPT, Claude, Copilot, Grok, and Gemini—provided incorrect answers to financial queries, the analysis showed.
Details
The study involved both free and paid AI models, including ChatGPT from OpenAI, Gemini from Google, Claude from Anthropic, and Copilot from Microsoft. As part of the test, researchers from Saturn—a firm that advises businesses on potential transformation through AI—asked the models more than 100 different questions on financial topics. The questions were repeated up to five times. As a result, a total of more than 10,000 questions were posed to 18 different AI models, the FT notes.
However, the results were disappointing. For example, an error in the explanation of UK pension tax rules made by Claude Haiku 4.5 (one of the free models released last year), could have resulted in a pension plan participant having to pay the British tax authority £17,500 ($23,400), the study found. In another instance—in response to a question about student loans—Claude “made up a rule,” stating that a graduate could stop making payments if they moved abroad.
Overall, when answering complex questions requiring more than one calculation, AI models made errors in 88% of cases on average, the researchers concluded. Moreover, some models provided incorrect answers to 99% of such complex questions. In particular, the neural networks’ answers contained calculation errors, failed to account for upcoming changes in taxation, or included hallucinations regarding the rules.
And although, according to Saturn’s analysis, paid neural networks provided more accurate answers than free ones, and newer AI models performed better than older ones—on average, the neural networks tested gave incorrect answers to financial questions 57% of the time, the researchers noted. The Claude Opus 5 model demonstrated the best results in “reasoning mode,” although even it made mistakes in 39% of cases.
“Millions of people trust AI models with their finances, but they receive incorrect answers that could cause them to lose money,” warned Amal Jolly, CEO of Saturn. That said, despite these findings, chatbots can be useful for budgeting and gathering information—for example, they can suggest which expense categories are consuming too much of a user’s budget, based on their income and spending patterns,” noted Sarah Coles, head of personal finance at the investment platform AJ Bell.
Context
The Saturn study was released a few days after Anthropic unveiled its “Claude for Financial Advisors” tool on September 14, which integrated a chatbot with the platforms of BlackRock, Charles Schwab, and Addepar to help financial advisors analyze portfolios and prepare for meetings. The release came shortly after OpenAI unveiled a similar industry-specific version of ChatGPT (the company combined its latest AI model with built-in data from providers such as LSEG, PitchBook, and Daloopa to assist investment bankers and analysts).
This article was AI-translated and verified by a human editor



