[{"data":1,"prerenderedAt":-1},["ShallowReactive",2],{"news-slug-mistral-leanstral-1-5":3,"news-related-c4acfb32-44d4-45a9-ac27-ec9fff4e1eca":36},{"id":4,"title":5,"summary":6,"content":6,"original_url":7,"source_id":8,"tags":9,"translations":23,"news_slug":29,"published_at":30,"created_at":31,"modified_at":32,"is_published":33,"publish_type":34,"image_url":13,"view_count":35},"c4acfb32-44d4-45a9-ac27-ec9fff4e1eca","Mistral 开源 Leanstral 1.5:6B 激活参数刷新形式化推理 SOTA","2026 年 7 月 2 日,Mistral AI 开源 **Leanstral 1.5**——面向 Lean 4 证明助手的专用 Agent,Apache-2.0 协议、免费可用。它采用 119B 总参 \u002F 6B 激活的 MoE 架构,把\"小激活、大容量\"范式带到了严肃形式化数学。\n\n成绩单很硬:miniF2F 验证 \u002F 测试双双 100% 饱和;PutnamBench 解 587\u002F672,比 Seed-Prover 1.5 high 多 7 题,单题成本仅约 4 美元(对方约 300 美元);FATE-H 87 题、FATE-X 34 题均为榜单第一。配套开源的 FLTEval 基准上,pass@8 推到 43.2,超越闭源 Opus 4.6 的 39.6,成本仅其 1\u002F7。\n\n训练管线走\"继续预训练 → SFT → **CISPO** 强化学习\"三段式,RL 同时跑在多轮 Lean 验证环境与代码 Agent 环境。最惊艳的是 test-time scaling:PutnamBench Pass@8 随 token 预算从 50k 到 4M 单调爬升,44 → 244 → 493 → 587;AVL 树 O(log n) 的证明跑了 2.7M token、22 轮上下文压缩,是这条曲线的极限样本。\n\n实战层面,对 57 个真实 Rust 仓库扫描,模型抓到 47 处可疑属性、11 个真实 bug,5 个是 GitHub 上从未报告的——包括 varinteger 库 zigzag 解码在 u64::MAX 时的整数溢出,fuzzing 都难以命中。\n\n评论:真正信号不在榜单,而是\"6B 激活 + 完整 RL 工具链 + Apache-2.0\"这套组合首次把严肃形式化验证做成了研究机构之外也能部署的工程基座。从今往后,数学推理、关键代码审计、形式化安全验证团队都有了一个不会被 API 费用卡脖子的开源底座。","https:\u002F\u002Fmistral.ai\u002Fnews\u002Fleanstral-1-5\u002F","2436174c-644b-4a65-9a98-e7a3b705569a",[10,14,17,20],{"id":11,"name":12,"slug":12,"description":13,"color":13},"120fa59a-ff6f-4537-9bf5-f818df636a0e","benchmark",null,{"id":15,"name":16,"slug":16,"description":13,"color":13},"e82b2d09-81b2-43d1-977e-e018443b3c14","coding-agent",{"id":18,"name":19,"slug":19,"description":13,"color":13},"01598627-1ea6-4b27-a5d8-874971571a71","llm",{"id":21,"name":22,"slug":22,"description":13,"color":13},"b9bd9039-fcdb-41a8-b85b-fc1587def2b9","open-source",[24],{"id":25,"lang":26,"title":27,"summary":28,"content":13},"53142b21-a424-4363-9d25-59b9b9aab7e7","en","Leanstral 1.5: Mistral's 6B-active formal reasoning SOTA","On July 2, 2026, Mistral AI open-sources **Leanstral 1.5** — a dedicated Agent for the Lean 4 proof assistant, under Apache-2.0 license, free to use. It uses a 119B-total \u002F 6B-activated MoE architecture, bringing the \"small activation, large capacity\" paradigm to serious formal mathematics. The scorecard is hard: miniF2F validation\u002Ftest both saturated at 100%; PutnamBench solved 587\u002F672, 7 more than Seed-Prover 1.5 high, with single-problem cost only about $4 (versus about $300); FATE-H 87 problems, FATE-X 34 problems both rank first on their leaderboards. On the companion open-source FLTEval benchmark, pass@8 climbs to 43.2, surpassing closed-source Opus 4.6's 39.6, at only 1\u002F7 the cost. The training pipeline takes the three-stage path \"continued pretraining → SFT → **CISPO** RL\", with RL running on both multi-turn Lean verification environments and code-Agent environments. The most amazing is test-time scaling: PutnamBench Pass@8 climbs monotonically with token budget from 50k to 4M, 44 → 244 → 493 → 587; the AVL tree O(log n) proof ran 2.7M tokens, 22 rounds of context compression, an extreme sample of this curve. On the practical level, scanning 57 real Rust repositories, the model caught 47 suspicious properties, 11 real bugs, 5 of which were never reported on GitHub — including integer overflow in the varinteger library's zigzag decoding at u64::MAX, hard to even hit with fuzzing. Commentary: the real signal isn't the leaderboard, but the \"6B activation + complete RL toolchain + Apache-2.0\" combination making serious formal verification an engineering base that can be deployed outside of research institutions for the first time. From now on, math reasoning, critical code auditing, and formal security verification teams all have an open-source base that won't be choked by API costs.","mistral-leanstral-1-5","2026-07-04T00:30:00Z","2026-07-04T00:08:16.575060Z","2026-08-19T02:08:40.142862Z",true,"agent",101,{"items":37},[38,43,48,53,58,63],{"id":39,"title":40,"news_slug":41,"published_at":42},"7ba15299-8bee-4039-8bc7-dbb58754b562","SWE-bench Science:最强 Claude Code 修科学代码也不及格,四类失败模式被拆解","swe-bench-science-benchmark","2026-08-21T13:00:00+00:00",{"id":44,"title":45,"news_slug":46,"published_at":47},"5b76d928-8896-4bf8-8edd-e195ecf0094a","Ornith-1.5自报跑分赢了Claude,独立复测翻车了","ornith-1-5-benchmark-reality-check","2026-08-20T17:10:00+00:00",{"id":49,"title":50,"news_slug":51,"published_at":52},"ff0bc92a-295a-4707-be8d-76115fe9eeee","PerceptionBench 出炉:16 个前沿多模态模型,视觉感知无一及格","moonshot-perceptionbench-atomic-perception","2026-08-26T13:15:00+00:00",{"id":54,"title":55,"news_slug":56,"published_at":57},"f9cf9f03-6aca-4d29-94d3-5c6acfeaf435","匿名模型 OX Alpha 短暂登顶 OpenRouter 编码榜:研究者推测底座指向智谱 GLM-5.x","ox-alpha-stealth-openrouter-glm-5-zhipu","2026-08-24T03:00:00+00:00",{"id":59,"title":60,"news_slug":61,"published_at":62},"5a274662-0c3f-492c-a0e8-a46c5a0be783","别再自己给自己打分了:Co-RL 让模型互相判卷,无标签 RL 追平有监督","co-rl-peer-reward-label-free-rl","2026-08-20T19:10:00+00:00",{"id":64,"title":65,"news_slug":66,"published_at":67},"89a79f9a-bfd2-4ebe-8f03-92fa74a3a34f","Ornith-1.5 开源：模型自己出题、自己搭考场，397B 到 9B 三档齐发","ornith-1-5-self-improvement-open-models","2026-08-20T13:30:00+00:00"]