[{"data":1,"prerenderedAt":-1},["ShallowReactive",2],{"news-slug-florence-sota-llm-dice-probabilistic-fail":3,"topics-all":33,"news-related-05f3d81d-bad5-433c-a45f-a4cc3af7fc09":52},{"id":4,"title":5,"summary":6,"content":6,"original_url":7,"source_id":8,"tags":9,"translations":20,"news_slug":26,"published_at":27,"created_at":28,"modified_at":29,"is_published":30,"publish_type":31,"image_url":13,"view_count":32},"05f3d81d-bad5-433c-a45f-a4cc3af7fc09","SOTA 大模型玩骰子也翻车：佛罗伦萨大学论文揭 LLM 概率推理\"靠题感不靠推理\"","佛罗伦萨大学团队在 arXiv 发布 2606.07515 论文，把 8 个 SOTA 大模型（16 种有\u002F无 CoT 的配置）拉去做\"形式可证、又能触发直觉偏差\"的离散概率推理题。结果是一道分水岭：标准题平均 0.96（16 个里 9 个超 0.99），反直觉变体直接掉到 0.59，最强 ChatGPT 5.4 Thinking 也只到 0.84。论文接着做三种\"降维打击\"：把题面措辞改写成同构但陌生的版本，准确率掉 20%；在 prompt 里植入由其他模型生成的\"看上去合理\"的错误答案，性能最高崩 34%，且没有模型免疫；最反常的是 Mistral Large 3，开 CoT 几乎无收益。结论很直白——今天的 LLM 不是概率推理者，而是\"训练语料里的概率题复读机\"。它们在标准题上的稳健性更多来自对题面模板的检索，而非对概率公理的内部验证；RLHF 阶段的\"讨好\"训练又把推理天花板锁死在\"题感\"上。这恰好解释了为什么最近 GRPO、On-Policy Distillation 等工作开始把纠错压力从结果层推向 rollout 层——纯靠题感的红利快要到头了。","https:\u002F\u002Farxiv.org\u002Fabs\u002F2606.07515","7437aeb9-930c-4866-a2e9-48003c1a792b",[10,14,17],{"id":11,"name":12,"slug":12,"description":13,"color":13},"120fa59a-ff6f-4537-9bf5-f818df636a0e","benchmark",null,{"id":15,"name":16,"slug":16,"description":13,"color":13},"0a93ec8e-ea39-4693-81de-563ca8c173f7","inference",{"id":18,"name":19,"slug":19,"description":13,"color":13},"01598627-1ea6-4b27-a5d8-874971571a71","llm",[21],{"id":22,"lang":23,"title":24,"summary":25,"content":13},"ca537f54-8fa0-4f25-b395-3babe3b99402","en","SOTA LLMs flunk dice: probabilistic reasoning by vibes","A team at the University of Florence published arXiv paper 2606.07515, putting 8 SOTA large models (16 configurations with\u002Fwithout CoT) through discrete probability reasoning problems that are \"formally provable and can trigger intuitive bias.\" The result is a watershed: the standard question averages 0.96 (9 of 16 exceed 0.99), and the counter-intuitive variants drop directly to 0.59, with the strongest ChatGPT 5.4 Thinking only reaching 0.84. The paper then conducts three \"dimensionality reduction attacks\": rewriting the question to an isomorphic but unfamiliar version drops accuracy by 20%; planting \"reasonable-looking\" wrong answers generated by other models in the prompt drops performance by up to 34%, with no model immune; the most unusual is Mistral Large 3, which gains almost nothing from turning on CoT. The conclusion is direct: today's LLMs are not probabilistic reasoners, but \"probability question repeaters\" in the training corpus. Their robustness on standard questions comes more from retrieving the question template than from internal verification of probability axioms; the \"pleasing\" training of the RLHF stage also locks the reasoning ceiling on \"question feel.\" This exactly explains why recent works like GRPO and On-Policy Distillation have started pushing the correction pressure from the result layer to the rollout layer — the pure-question-feel dividend is about to end.","florence-sota-llm-dice-probabilistic-fail","2026-06-09T02:25:00Z","2026-06-09T02:25:15.732170Z","2026-08-19T02:08:40.142862Z",true,"agent",145,[34,43],{"slug":35,"tag_slug":35,"title_zh":36,"title_en":37,"intro_zh":38,"intro_en":39,"id":40,"is_active":30,"created_at":41,"modified_at":42},"ai-for-science","AI for Science 2026：从 UniPert 到 GPT-Rosalind 的硬核进化","AI for Science 2026: from UniPert to GPT-Rosalind","生命科学、化学材料、物理世界模型——AI 正在从\"语言工具\"变成\"实验伙伴\"。本专题收录 AI 在三大科学方向的关键节点：UniPert 统一基因与化学扰动空间、GPT-Rosalind 端到端生命科学推理、达摩院 AI 智能体 28 小时找到 4 种超导新材料、Anthropic Claude Science 把工作台做成标准品。","From language tool to lab partner — AI is reshaping life sciences, chemistry\u002Fmaterials, and physical world models. This topic covers the key milestones: UniPert unifying genetic-chemical perturbation spaces, GPT-Rosalind's end-to-end life-sciences reasoning, DAMO's AI agent discovering 4 superconducting materials in 28 hours, and Anthropic's Claude Science workbench going mainstream.","988a4300-5fab-41c4-b5d8-63711a2dc757","2026-09-10T01:34:15.296649Z","2026-09-10T01:34:15.296663Z",{"slug":44,"tag_slug":44,"title_zh":45,"title_en":46,"intro_zh":47,"intro_en":48,"id":49,"is_active":30,"created_at":50,"modified_at":51},"h3-series","MiniMax H3 系列：从开源权重到 35 倍吞吐","MiniMax H3 Series: from open weights to 35x throughput","MiniMax H3 自 2026 年 8 月开源以来节奏密集：官方把生成、参考与编辑收回一个模型；ComfyUI 当天压进 RTX 3060；摩尔线程 3 小时完成国产 GPU 适配；fal 后训练版把吞吐拉到 35 倍；FastH3 蒸馏再砍推理成本。本专题持续追踪 H3 的发布—开源—蒸馏—部署全链路。","Since MiniMax open-sourced H3 in August 2026 the pace has been relentless: one unified omni-modal model, same-day ComfyUI support down to an RTX 3060, a 3-hour Day-0 port to Moore Threads GPUs, fal's post-trained H3 Max at 35x throughput, and FastH3 distillation cutting inference cost further. This topic tracks the full H3 chain — release, open weights, distillation, deployment.","83ef0daa-3c31-4cb3-86ed-e5ee58654d5f","2026-09-08T07:33:19.942193Z","2026-09-08T07:33:19.942209Z",{"items":53},[54,59,64,69,74,79],{"id":55,"title":56,"news_slug":57,"published_at":58},"13ed0c94-9bea-440c-93c4-771c38ba9834","sink 消失后，KV 驱逐改看 value 几何","valuediff-value-geometric-kv-eviction","2026-09-22T17:05:00+00:00",{"id":60,"title":61,"news_slug":62,"published_at":63},"e2eb81dc-3112-411a-94b0-b5061a12be78","AdvancedMathBench 把数学证明拉进博士级:GPT-5.5-xhigh 仍只 75.8","advanced-math-bench-phd-level","2026-07-14T16:15:00+00:00",{"id":65,"title":66,"news_slug":67,"published_at":68},"fc1e888d-4bee-4633-86f4-edc76aa48161","BlockSearch 把语言模型变成「上下文检索器」：0.6B 在百万 token 上打平向量检索","blocksearch-context-retriever","2026-07-03T06:25:00+00:00",{"id":70,"title":71,"news_slug":72,"published_at":73},"7cacddc6-fa02-4de9-84a2-c3320e225571","因果归因剪枝 CAP：让 LLM 推理能力不再随稀疏化而流失","cap-causal-attribution-pruning-arc-61pct","2026-06-20T22:14:08.915874+00:00",{"id":75,"title":76,"news_slug":77,"published_at":78},"2d866988-f04f-42e8-85e1-8a49069f9222","MaxProof测试时缩放：MiniMax M3拿下IMO 2025\u002F USAMO 2026双金","minimax-m3-maxproof-imo-2025-usamo-gold","2026-06-14T12:00:00+00:00",{"id":80,"title":81,"news_slug":82,"published_at":83},"7e7d7d95-a592-4ee7-a43f-c3d109c58405","LLM 长上下文「有效容量」被高估了：12K 词也能撑爆，密度才是隐藏分水岭","dense-contexts-12k-needle-density-mattr","2026-06-08T08:00:00+00:00"]