Ask ChatGPT whether you should pay off your mortgage early or fund your pension instead, and there is a better-than-even chance it gets the answer wrong. That is the core finding of "Artificial Authority: Should you trust AI to deliver financial advice?", a September 2026 report by fintech firm Saturn, which tested 18 mainstream AI models and found them wrong on 57% of financial questions - with the error rate climbing to 88% on the hardest ones.
121 real questions, ten thousand answers
Saturn picked 121 real financial questions people search for every week, covering debt, student loans, mortgages, pensions, tax and savings, and put them to 18 models including ChatGPT, Claude, Copilot, Grok and Gemini. Each question was repeated five times to check consistency, producing more than 10,000 answers. On average, 57% were wrong - an accuracy rate of just 43%. The harder the question, the worse it got: on the toughest multi-step problems the average error rate hit 88%, and some models were wrong 99% of the time. These are not trick questions; they are things like "what happens if I miss a student loan payment" that millions of people type into search boxes every week.
The failure modes matter more than the rate
According to a Solidot write-up of the report, the mistakes are not just arithmetic: models missed imminent tax-policy changes and flat-out fabricated rules - textbook hallucination. The study also found paid models more accurate than free ones, and newer models better than older ones; even the best reasoning-mode model still got 39% of answers wrong.
The demand side: 57% act without checking
Saturn's data covers the supply side. A PensionBee survey of 1,000 US adults who use AI chatbots for personal finance fills in the demand side: 57% said they would act on a chatbot's money advice without verifying it first; 23% have already received wrong financial information from a chatbot, and 5% only found out after acting on it. The generational split is stark: 66% of Gen Z would let an AI act autonomously on their behalf, financial decisions included, versus 47% of Baby Boomers - and the cohort most comfortable handing over the wheel is often the one with the least savings and the least room to absorb a wrong call. Fortune, citing Gallup, reports that one in five Americans already use AI for financial advice while seven in ten say they don't trust it. Both things are true at once.
Regulators lag, industry rushes ahead
The UK's Financial Conduct Authority flagged exactly this gap in its Mills Review this year, warning that AI tools increasingly blur the line between guidance and regulated advice, and urging the government to consider moving the regulatory boundary before real harm piles up. Industry is not waiting: Worldline, ING and Mastercard ran Europe's first end-to-end agentic payment in June, with an AI agent executing a real payment inside the banking system; Santander and Mastercard have run a similar live test in a regulated environment. Independent research shows these same models get the underlying financial math wrong more often than right, while card companies race to put AI agents in charge of spending decisions. Saturn CEO Amal Jolly's warning is blunt: the low quality of financial advice from mainstream AI models risks widespread consumer harm.
So what
A licensed adviser who gets basic questions wrong 57% of the time loses their license; a model that does the same gets a version bump. None of this makes AI useless for money questions - it is fast, free, and can explain the general shape of a mortgage refinance well enough to get you started. The trouble starts when someone treats a fluent answer as a correct one, then hands over a decision that is hard to undo. While model error rates are over half, regulation is blank, and users skip verification, the right way to use AI financial advice is: ask, then get a second human to check.
Sources: Saturn report; Startup Fortune coverage.