[{"data":1,"prerenderedAt":-1},["ShallowReactive",2],{"news-slug-saturn-ai-financial-advice-error-rate":3,"topics-all":38,"news-related-0da59714-49b8-4ea4-a74e-fbac3e8c532f":57},{"id":4,"title":5,"summary":6,"content":7,"original_url":8,"source_id":9,"tags":10,"translations":24,"news_slug":31,"published_at":32,"created_at":33,"modified_at":34,"is_published":35,"publish_type":36,"image_url":14,"view_count":37},"0da59714-49b8-4ea4-a74e-fbac3e8c532f","实测18个主流模型:财务问答平均57%答错,难题88%","金融科技公司Saturn用121道真实财务问题实测ChatGPT、Claude等18个主流AI模型,每题重复5次、累计逾万次作答,平均57%给出错误答案,最难题目出错率达88%,报告呼吁给AI理财问答先建护栏。","把「提前还房贷还是多缴养老金」这类问题抛给 ChatGPT，得到正确答案的概率不到一半。金融科技公司 Saturn 今年 9 月发布的报告《Artificial Authority: Should you trust AI to deliver financial advice?》实测了 18 个主流 AI 模型：平均 57% 的财务问题答错，难题出错率冲到 88%。\n\n## 121 道真实财务题，一万次作答\n\nSaturn 团队挑了 121 道普通人每周都在搜的真实财务问题，覆盖债务、学生贷款、房贷、养老金、税务和储蓄，扔给 ChatGPT、Claude、Copilot、Grok、Gemini 等 18 个模型。每道题重复问 5 次以检验一致性，累计作答超过 10,000 次。结果是平均 57% 的回答是错的，准确率只有 43%。题目越难越不可靠：在最难的多步问题上，平均出错率 88%，个别模型 99% 都是错的。这些不是陷阱题，而是「错过一次学生贷款还款会怎样」「该不该提前还房贷」这种每周都有人往搜索框里敲的问题。\n\n## 错误类型比错误率更扎心\n\n据 Solidot 对报告的编译转述，错误不只是算错数——有遗漏即将实施的税收政策变更的，也有凭空捏造规则的典型幻觉。研究还发现付费模型比免费模型准确，较新模型优于较旧型号；即便表现最好的推理模式模型，错误率仍有 39%。\n\n## 用户端：57% 的人不核对就直接照做\n\nSaturn 的数据只覆盖供给侧，PensionBee 对 1,000 名美国成年用户的调查补上了需求侧：57% 的受访者会不做任何独立核验、直接按 AI 建议操作；23% 已从聊天机器人收到过错误的财务信息，其中 5% 是照做之后才发现的。代际差异更明显：66% 的 Z 世代愿意让 AI 自主替自己处理包括财务在内的决策，婴儿潮一代只有 47%——而最愿意交出控制权的群体，恰恰是储蓄最少、最经不起错误决策的群体。Fortune 援引 Gallup 的调查称，五分之一的美国人已在用 AI 做财务建议，同时七成人表示不信任它，两件事同时为真。\n\n## 监管没跟上，产业已经在冲\n\n英国金融行为监管局（FCA）今年在 Mills Review 中专门警告：AI 工具正在模糊「信息指引」与「持牌建议」的边界，应在真实伤害堆积前移动监管红线。但产业没有等监管——Worldline、ING 和 Mastercard 今年 6 月在欧洲跑通了首笔端到端 agentic 支付，由 AI 智能体在银行系统内真实执行；Santander 和 Mastercard 也在受监管环境做了类似实测。一边是独立研究证明这些模型财务数学错多对少，一边是卡组织竞相把 AI agent 推向消费决策一线。Saturn CEO Amal Jolly 的警告很直白：主流 AI 模型财务建议的低质量，可能导致大范围的消费者伤害。\n\n## 所以呢\n\n持牌顾问基础问题答错 57% 会丢执照；模型做同样的事，得到的只是一次版本号升级。这不是说 AI 对财务问题毫无用处——它免费、快，能把再融资的大致框架讲清楚，适合当入门导读。真正的风险是把「流利」当「正确」，然后交出一个难以撤销的决策：什么时候退休、借多少钱。在模型错误率过半、监管空白、用户不核验三件事同时成立的当下，AI 财务建议的正确用法是：问完之后，找第二个人核对。\n\n参考：[Saturn 报告原文](https:\u002F\u002Fwww.saturnos.com\u002Freport\u002Fartificial-authority)；[Startup Fortune 报道](https:\u002F\u002Fstartupfortune.com\u002Fai-chatbots-get-financial-questions-wrong-more-than-half-the-time-study-finds\u002F)。","https:\u002F\u002Fwww.saturnos.com\u002Freport\u002Fartificial-authority","7c43bc02-0c7c-4cea-a471-b19bfa7fef06",[11,15,18,21],{"id":12,"name":13,"slug":13,"description":14,"color":14},"1fcfaaf2-67de-43d3-9e35-5784852fec60","ai-safety",null,{"id":16,"name":17,"slug":17,"description":14,"color":14},"40269b40-7942-4650-9672-ed2e6524d37a","ai-technology",{"id":19,"name":20,"slug":20,"description":14,"color":14},"120fa59a-ff6f-4537-9bf5-f818df636a0e","benchmark",{"id":22,"name":23,"slug":23,"description":14,"color":14},"01598627-1ea6-4b27-a5d8-874971571a71","llm",[25],{"id":26,"lang":27,"title":28,"summary":29,"content":30},"af278cb5-f117-40a5-b420-b28b48261cf3","en","Saturn tested 18 AI models: 57% wrong on financial advice","Saturn tested 18 AI models on 121 financial questions: 57% of answers were wrong, 88% on the hardest - yet users act on AI advice without checking.","Ask ChatGPT whether you should pay off your mortgage early or fund your pension instead, and there is a better-than-even chance it gets the answer wrong. That is the core finding of \"Artificial Authority: Should you trust AI to deliver financial advice?\", a September 2026 report by fintech firm Saturn, which tested 18 mainstream AI models and found them wrong on 57% of financial questions - with the error rate climbing to 88% on the hardest ones.\n\n## 121 real questions, ten thousand answers\n\nSaturn picked 121 real financial questions people search for every week, covering debt, student loans, mortgages, pensions, tax and savings, and put them to 18 models including ChatGPT, Claude, Copilot, Grok and Gemini. Each question was repeated five times to check consistency, producing more than 10,000 answers. On average, 57% were wrong - an accuracy rate of just 43%. The harder the question, the worse it got: on the toughest multi-step problems the average error rate hit 88%, and some models were wrong 99% of the time. These are not trick questions; they are things like \"what happens if I miss a student loan payment\" that millions of people type into search boxes every week.\n\n## The failure modes matter more than the rate\n\nAccording to a Solidot write-up of the report, the mistakes are not just arithmetic: models missed imminent tax-policy changes and flat-out fabricated rules - textbook hallucination. The study also found paid models more accurate than free ones, and newer models better than older ones; even the best reasoning-mode model still got 39% of answers wrong.\n\n## The demand side: 57% act without checking\n\nSaturn's data covers the supply side. A PensionBee survey of 1,000 US adults who use AI chatbots for personal finance fills in the demand side: 57% said they would act on a chatbot's money advice without verifying it first; 23% have already received wrong financial information from a chatbot, and 5% only found out after acting on it. The generational split is stark: 66% of Gen Z would let an AI act autonomously on their behalf, financial decisions included, versus 47% of Baby Boomers - and the cohort most comfortable handing over the wheel is often the one with the least savings and the least room to absorb a wrong call. Fortune, citing Gallup, reports that one in five Americans already use AI for financial advice while seven in ten say they don't trust it. Both things are true at once.\n\n## Regulators lag, industry rushes ahead\n\nThe UK's Financial Conduct Authority flagged exactly this gap in its Mills Review this year, warning that AI tools increasingly blur the line between guidance and regulated advice, and urging the government to consider moving the regulatory boundary before real harm piles up. Industry is not waiting: Worldline, ING and Mastercard ran Europe's first end-to-end agentic payment in June, with an AI agent executing a real payment inside the banking system; Santander and Mastercard have run a similar live test in a regulated environment. Independent research shows these same models get the underlying financial math wrong more often than right, while card companies race to put AI agents in charge of spending decisions. Saturn CEO Amal Jolly's warning is blunt: the low quality of financial advice from mainstream AI models risks widespread consumer harm.\n\n## So what\n\nA licensed adviser who gets basic questions wrong 57% of the time loses their license; a model that does the same gets a version bump. None of this makes AI useless for money questions - it is fast, free, and can explain the general shape of a mortgage refinance well enough to get you started. The trouble starts when someone treats a fluent answer as a correct one, then hands over a decision that is hard to undo. While model error rates are over half, regulation is blank, and users skip verification, the right way to use AI financial advice is: ask, then get a second human to check.\n\nSources: [Saturn report](https:\u002F\u002Fwww.saturnos.com\u002Freport\u002Fartificial-authority); [Startup Fortune coverage](https:\u002F\u002Fstartupfortune.com\u002Fai-chatbots-get-financial-questions-wrong-more-than-half-the-time-study-finds\u002F).","saturn-ai-financial-advice-error-rate","2026-09-21T17:30:00Z","2026-09-21T17:11:06.936706Z","2026-09-21T17:11:06.936719Z",true,"agent",1,[39,48],{"slug":40,"tag_slug":40,"title_zh":41,"title_en":42,"intro_zh":43,"intro_en":44,"id":45,"is_active":35,"created_at":46,"modified_at":47},"ai-for-science","AI for Science 2026：从 UniPert 到 GPT-Rosalind 的硬核进化","AI for Science 2026: from UniPert to GPT-Rosalind","生命科学、化学材料、物理世界模型——AI 正在从\"语言工具\"变成\"实验伙伴\"。本专题收录 AI 在三大科学方向的关键节点：UniPert 统一基因与化学扰动空间、GPT-Rosalind 端到端生命科学推理、达摩院 AI 智能体 28 小时找到 4 种超导新材料、Anthropic Claude Science 把工作台做成标准品。","From language tool to lab partner — AI is reshaping life sciences, chemistry\u002Fmaterials, and physical world models. This topic covers the key milestones: UniPert unifying genetic-chemical perturbation spaces, GPT-Rosalind's end-to-end life-sciences reasoning, DAMO's AI agent discovering 4 superconducting materials in 28 hours, and Anthropic's Claude Science workbench going mainstream.","988a4300-5fab-41c4-b5d8-63711a2dc757","2026-09-10T01:34:15.296649Z","2026-09-10T01:34:15.296663Z",{"slug":49,"tag_slug":49,"title_zh":50,"title_en":51,"intro_zh":52,"intro_en":53,"id":54,"is_active":35,"created_at":55,"modified_at":56},"h3-series","MiniMax H3 系列：从开源权重到 35 倍吞吐","MiniMax H3 Series: from open weights to 35x throughput","MiniMax H3 自 2026 年 8 月开源以来节奏密集：官方把生成、参考与编辑收回一个模型；ComfyUI 当天压进 RTX 3060；摩尔线程 3 小时完成国产 GPU 适配；fal 后训练版把吞吐拉到 35 倍；FastH3 蒸馏再砍推理成本。本专题持续追踪 H3 的发布—开源—蒸馏—部署全链路。","Since MiniMax open-sourced H3 in August 2026 the pace has been relentless: one unified omni-modal model, same-day ComfyUI support down to an RTX 3060, a 3-hour Day-0 port to Moore Threads GPUs, fal's post-trained H3 Max at 35x throughput, and FastH3 distillation cutting inference cost further. This topic tracks the full H3 chain — release, open weights, distillation, deployment.","83ef0daa-3c31-4cb3-86ed-e5ee58654d5f","2026-09-08T07:33:19.942193Z","2026-09-08T07:33:19.942209Z",{"items":58},[59,64,69,74,79,84],{"id":60,"title":61,"news_slug":62,"published_at":63},"b05de01b-89ca-499b-b130-e55162e651f5","SCOPE：让大模型学会选择性信任，而不是把上下文一概拒绝","scope-selective-trust-context-dpo","2026-08-06T17:59:58+00:00",{"id":65,"title":66,"news_slug":67,"published_at":68},"f8207ff2-88ad-4e11-a87c-8350ffd42c01","推理训练在悄悄「偷走」模型对齐：arXiv 新论文六大维度系统审计","reasoning-alignment-audit-6-dim-2606-11046","2026-06-10T12:15:00+00:00",{"id":70,"title":71,"news_slug":72,"published_at":73},"92eaa312-6506-4314-8fa5-f171ce0f8ea2","伯克利研究撕开AI评测遮羞布：所有主流Agent基准均可被免解题刷到满分","berkeley-trustworthy-benchmarks-8-agent-gamed","2026-05-09T19:10:00+00:00",{"id":75,"title":76,"news_slug":77,"published_at":78},"3bcb0e1d-99ea-4fae-9bdd-b6b625aabf10","代码 agent 8 成都在骗你:12 模型实测揭晓","overclaimbench-llm-agents","2026-09-21T07:00:00+00:00",{"id":80,"title":81,"news_slug":82,"published_at":83},"58ed753e-ad6d-4aac-95f4-36bf217e169c","把 10 万条人类视频变成机器人教材:RoboTok 检索 mAP 提升约 50 倍,hard 任务 79.3% 对 19.5%","robotok-retrieval-benchmark-reread","2026-09-06T21:11:25+00:00",{"id":85,"title":86,"news_slug":87,"published_at":88},"47f12a0d-8559-473f-97be-dc12966bd4ff","DeepMind 双盲评测：Gemini 权重和考题锁进同一个加密飞地","deepmind-gemini-double-blind-eval","2026-08-29T15:05:00+00:00"]