Story resurfaces on Hacker News and Mastodon
2 Sep 21 12:28 AM · 5d ago · 1 article · 2 posts · 3 sources · development 2 of 2
Two days after initial publication, the FT article gained renewed attention, appearing on Hacker News's front page and being shared on Mastodon, though with minimal accompanying discussion in the available material.
Financial Times Publisher of the report
The whole story articlesposts the bright band is this development · numbered dots are the others · click one to jump
Reported in the same hours no headline names this development itself — these 1 claim were published in its stretch
-
first by HN Frontpage, 5d ago · also Tom's Guide
1 more headline
What people said 13 voices · best of 15 · verbatim
-
So I downloaded that report which of course doesn't contain the most relevant information (the questions) but it contains some examples of wrong answers.I fed the first question to Grok (which they claimed they tested as well) and it answered it correctly in detail.I repeated it with another one - again correct answer. I then selected the question…
-
R
Financial Misinformation, Tilly Norwood, Anti-Data Center Activism, More: ResearchBuzz AI Update, September 20, 2026 AI PROBLEMS Financial Times: AI chatbots give wrong answers to financial queries ‘most of the time’... http:// researchbuzz.me/2026/09/20/fin…
-
This will be the same story for every industry again and again. AI is not good in x=Finance because models were not RL trained heavily on x=Finance capabilities yet. This is only because Big Labs have finance benchmarks lower in their priority list. Their first priority was solving programming because that gives the best leverage at this stage. As…
-
Y
"Popular AI chatbots give incorrect or incomplete answers to financial questions more often than they get them right. . . .the average accuracy rate was just 43%, meaning AI chatbots made mistakes 57% of the time. . ." https://www. investmentnews.com/fintech/ai- chatbots-give-wrong-financial-answers-most-of-the-time-study-finds/268267
-
The problem with LLMs in finance is the same as it is in writing, design, and many other disciplines: it isn’t code.Code objectively does what it‘s intended to do or it doesn’t (and passes certain tests or not) which gives coding agents an indication on whether their solution is adequate.This is much harder in almost any other discipline.Pass-fail…
-
Article is just a vague summary of https://www.saturnos.com/report/artificial-authorityAnecdotally, current models seem to be decent at general personal finance principles - certainly better than the majority of personal finance education that people get exposed to unless they seek it out and read a variety of books and sources. But I wouldn't…
-
These models do pretty well in benchmarks and real world so I'm highly suspicious of this article. Further more, in the original report, the examples of bad answers are from Haiku - at least 7 out of 10. Anyone who knows anything about LLMs know that haiku shouldn't be used for anything pretty much.There's no reproducible set either. I'm not gonna…
-
And they hallucinate errors in the millions and struggle with financial data that is in a layout that isn’t in the training data. Ie balance sheet etc.Been trying to add more AI to my workflow but it just doesn’t work (yet) - not in the same way as vibe coding doesThe technical references lookups work though. Looking up regulations etc
-
I think this depends a lot on what context you give it. I've had solid answers on financial stuff when I give it the right info. But, I wouldn't trust it to guess the missing pieces. I'd be curious how much data or context the model's had in the test.
-
Given most financial advisors tend to vend out suboptimal advice and steer customers in favour of products they receive a kickback for, I'm happy to be accepting of an unbiased LLM that's trained on bogleheads.org.
-
Each year I have Claude do my taxes (which are complex) and compare them to those of our tax consultant.Each year it's exactly the same.I guess I'm getting really lucky?
-
Single shot or with reasoning enabled? My experience is that reasoning dramatically reduces hallucinations and improves output quality. I don't trust models without it.
-
Is this just because LLMs can't do math directly? If so, they can certainly write scripts to do math though and those will be a lot more reliable for queries.
All 2 developments of FT Report: AI Chatbots Wrong on Financial Queries "Most of… →
BlueskyMastodonHacker NewsNewswires