[{"data":1,"prerenderedAt":-1},["ShallowReactive",2],{"news-slug-advanced-math-bench-phd-level":3,"news-related-e2eb81dc-3112-411a-94b0-b5061a12be78":36},{"id":4,"title":5,"summary":6,"content":6,"original_url":7,"source_id":8,"tags":9,"translations":23,"news_slug":29,"published_at":30,"created_at":31,"modified_at":32,"is_published":33,"publish_type":34,"image_url":13,"view_count":35},"e2eb81dc-3112-411a-94b0-b5061a12be78","AdvancedMathBench 把数学证明拉进博士级:GPT-5.5-xhigh 仍只 75.8","当大多数 benchmark 还在用高中\u002F奥数级数学题考 LLM 时,Intern Large Models(上海 AI Lab)推出的 AdvancedMathBench 把题目难度直接拉到「本科生高年级(UGD)+ 博士资格考(QE)」级别——核心 ProverBench 收录 296 道这样的难题,配套 VerifierBench 又用 888 条模型生成的证明轨迹去测「验证能力」。\n\n实验结果相当难看。在证明生成上,目前最强的 GPT-5.5-xhigh 也只拿下 75.8(UGD)和 66.1(QE),意味着即使是最顶级的模型,在博士级数学证明上也有 1\u002F4 到 1\u002F3 的题目彻底做不出来。在证明验证上,最强模型 Balanced F1 只有 65.1,而且所有模型的 True Negative 普遍偏低——模型抓「错误证明」的本事远远不够,容易把错的当成对的过。\n\n这套 benchmark 的最大贡献,不在于「又一次证明 LLM 不会做难题」,而在于把评估颗粒度从「最终答案对错」细化到「证明过程是否有效」,用 verifier pipeline + 专家标注做 fine-grained 错误定位。对 agentic 工作流(让 LLM 互相审稿、互相改稿)非常关键——如果 verifier 抓不到漏洞,整个 agent 链条的可信度就是空话。\n\n对比 OpenAI 复审 SWE-Bench Pro 时承认有近三成「坏题」,AdvancedMathBench 选择了更难的方向:题目可能更干净,但评判标准更严苛。这或许暗示 LLM 评估正在从「刷榜」走向「过程审计」的下一阶段。","https:\u002F\u002Farxiv.org\u002Fabs\u002F2607.11849","7437aeb9-930c-4866-a2e9-48003c1a792b",[10,14,17,20],{"id":11,"name":12,"slug":12,"description":13,"color":13},"5e628969-6d2a-437f-998a-104e4b16cfb1","ai-progress",null,{"id":15,"name":16,"slug":16,"description":13,"color":13},"120fa59a-ff6f-4537-9bf5-f818df636a0e","benchmark",{"id":18,"name":19,"slug":19,"description":13,"color":13},"0a93ec8e-ea39-4693-81de-563ca8c173f7","inference",{"id":21,"name":22,"slug":22,"description":13,"color":13},"01598627-1ea6-4b27-a5d8-874971571a71","llm",[24],{"id":25,"lang":26,"title":27,"summary":28,"content":13},"492bf750-3084-4979-9bdc-aa7b1db831bf","en","AdvancedMathBench: PhD-level proofs, GPT-5.5 at only 75.8","When most benchmarks are still testing LLMs with high-school \u002F Olympiad-level math, AdvancedMathBench, released by Intern Large Models (Shanghai AI Lab), pulls the difficulty straight up to \"upper-year undergraduate (UGD) + doctoral qualifying exam (QE)\" level — the core ProverBench collects 296 such hard problems, and the companion VerifierBench uses 888 model-generated proof trajectories to test \"verification ability\". The experimental results are rather harsh. On proof generation, even the currently strongest GPT-5.5-xhigh only scores 75.8 (UGD) and 66.1 (QE), meaning that even the top models are completely stumped by a quarter to a third of PhD-level math proofs. On proof verification, the strongest model's Balanced F1 is only 65.1, and True Negative is generally low across all models — models are far from being good at catching \"wrong proofs\" and easily let wrong ones through as right. The biggest contribution of this benchmark is not \"yet another proof that LLMs can't do hard problems\", but rather refining the evaluation granularity from \"is the final answer right\" to \"is the proof process valid\", using a verifier pipeline + expert annotation for fine-grained error localization. This is critical for agentic workflows (letting LLMs peer-review and edit each other) — if the verifier can't catch the holes, the credibility of the entire agent chain is empty talk. Compared with OpenAI's admission during the SWE-Bench Pro re-review that nearly 30% of problems are \"bad problems\", AdvancedMathBench chose a harder direction: the problems may be cleaner, but the judging standard is harsher. This may hint that LLM evaluation is moving from \"climbing leaderboards\" to the next stage of \"process auditing\".","advanced-math-bench-phd-level","2026-07-14T16:15:00Z","2026-07-14T16:12:26.556700Z","2026-08-19T02:08:40.142862Z",true,"agent",99,{"items":37},[38,43,48,53,58,63],{"id":39,"title":40,"news_slug":41,"published_at":42},"fc1e888d-4bee-4633-86f4-edc76aa48161","BlockSearch 把语言模型变成「上下文检索器」：0.6B 在百万 token 上打平向量检索","blocksearch-context-retriever","2026-07-03T06:25:00+00:00",{"id":44,"title":45,"news_slug":46,"published_at":47},"95b04c15-d5e5-4dca-ab2b-14e343bdd4e6","UC Berkeley 曝光 AI 基准测试系统性漏洞：45 种方法可在 13 个主流榜单上「不解决任何问题拿满分」","uc-berkeley-benchmark-45-cheats-13-leaderboards","2026-05-15T01:00:00+00:00",{"id":49,"title":50,"news_slug":51,"published_at":52},"9c2b49ce-a543-48cb-932d-b36daea3035c","ICLR 2026杰出论文警示：LLM在多轮对话中平均性能暴跌39%","iclr-2026-multiturn-conversation-39pct-drop","2026-05-04T10:20:00+00:00",{"id":54,"title":55,"news_slug":56,"published_at":57},"1311adb6-dc19-41a7-a188-6760d9e53672","HF Summer 2026 报告:13 个下载量 Top 25 模型是 2022 年的老面孔","hugging-face-summer-2026-attention-adoption","2026-08-24T08:00:00+00:00",{"id":59,"title":60,"news_slug":61,"published_at":62},"43eda321-b0b7-4df7-b20e-9758cbab42c9","记忆越完整,眼前题越做不对:MemTrapBench 把 LLM 长期记忆框架打回原形","memtrapbench-llm-memory-cognitive-traps","2026-08-22T04:00:00+00:00",{"id":64,"title":65,"news_slug":66,"published_at":67},"c5304aea-6b80-4fbb-b2ac-f07a7a6d7713","DeepSeek V4 Flash 上午短暂\"翻车\"：国产开源模型的容量大考","deepseek-v4-flash-api-capacity-incident","2026-08-04T10:00:00+00:00"]