FT Report: AI Chatbots Wrong on Financial Queries "Most of the Time"
A Financial Times investigation finds leading AI chatbots frequently answer financial questions incorrectly, sparking wide social sharing.
What to know
- A Financial Times report concludes AI chatbots answer financial queries incorrectly 'most of the time'.
- The finding spread rapidly across social platforms (Bluesky, Mastodon) and aggregators (Hacker News) within about 48 hours.
- The underlying FT article's full methodology and specific chatbots tested are not detailed in the available material.
Financial Times Publisher of the report
How it unfolded 2 developments, newest first · click a bar or a number to jump articlesposts
-
2
Story resurfaces on Hacker News and Mastodon
Two days after initial publication, the FT article gained renewed attention, appearing on Hacker News's front page and being shared on Mastodon, though with minimal accompanying discussion in the available material.
-
So I downloaded that report which of course doesn't contain the most relevant information (the questions) but it contains some examples of wrong answers.I fed the first question to Grok (which they claimed they tested as well) and it answered it correctly in detail.I repeated it with another one - again correct answer. I then selected the question…
2 more of the top 3 · 15 posts in this stretch
-
R
Financial Misinformation, Tilly Norwood, Anti-Data Center Activism, More: ResearchBuzz AI Update, September 20, 2026 AI PROBLEMS Financial Times: AI chatbots give wrong answers to financial queries ‘most of the time’... http:// researchbuzz.me/2026/09/20/fin…
-
This will be the same story for every industry again and again. AI is not good in x=Finance because models were not RL trained heavily on x=Finance capabilities yet. This is only because Big Labs have finance benchmarks lower in their priority list. Their first priority was solving programming because that gives the best leverage at this stage. As…
-
- 1 day quiet
-
1
Financial Times reports AI chatbots often wrong on finance questions
The Financial Times published findings that AI chatbots give wrong answers to financial queries most of the time, and shared the story on its Bluesky account.
“AI chatbots give wrong answers to financial queries 'most of the time'”
— Financial Times -
first by HN Frontpage, 6d ago · also Tom's Guide
1 more headline
-
What people are saying 10 voices from 2 sites · best of 15 · verbatim
- Sep 22
-
Y
"Popular AI chatbots give incorrect or incomplete answers to financial questions more often than they get them right. . . .the average accuracy rate was just 43%, meaning AI chatbots made mistakes 57% of the time. . ." https://www. investmentnews.com/fintech/ai- chatbots-give-wrong-financial-answers-most-of-the-time-study-finds/268267
- Sep 21
-
Is this just because LLMs can't do math directly? If so, they can certainly write scripts to do math though and those will be a lot more reliable for queries.
-
Each year I have Claude do my taxes (which are complex) and compare them to those of our tax consultant.Each year it's exactly the same.I guess I'm getting really lucky?
-
The problem with LLMs in finance is the same as it is in writing, design, and many other disciplines: it isn’t code.Code objectively does what it‘s intended to do or it doesn’t (and passes certain tests or not) which gives coding agents an indication on whether their solution is adequate.This is much harder in almost any other discipline.Pass-fail…
-
I think this depends a lot on what context you give it. I've had solid answers on financial stuff when I give it the right info. But, I wouldn't trust it to guess the missing pieces. I'd be curious how much data or context the model's had in the test.
-
And they hallucinate errors in the millions and struggle with financial data that is in a layout that isn’t in the training data. Ie balance sheet etc.Been trying to add more AI to my workflow but it just doesn’t work (yet) - not in the same way as vibe coding doesThe technical references lookups work though. Looking up regulations etc
-
Given most financial advisors tend to vend out suboptimal advice and steer customers in favour of products they receive a kickback for, I'm happy to be accepting of an unbiased LLM that's trained on bogleheads.org.
-
These models do pretty well in benchmarks and real world so I'm highly suspicious of this article. Further more, in the original report, the examples of bad answers are from Haiku - at least 7 out of 10. Anyone who knows anything about LLMs know that haiku shouldn't be used for anything pretty much.There's no reproducible set either. I'm not gonna…
-
Single shot or with reasoning enabled? My experience is that reasoning dramatically reduces hallucinations and improves output quality. I don't trust models without it.
-
Article is just a vague summary of https://www.saturnos.com/report/artificial-authorityAnecdotally, current models seem to be decent at general personal finance principles - certainly better than the majority of personal finance education that people get exposed to unless they seek it out and read a variety of books and sources. But I wouldn't…