[{"data":1,"prerenderedAt":-1},["ShallowReactive",2],{"news-slug-flowbalance-verifier-gated-self-distillation":3,"topics-all":38,"news-related-21c7dec1-f68e-4641-974c-ae2bce87393e":57},{"id":4,"title":5,"summary":6,"content":7,"original_url":8,"source_id":9,"tags":10,"translations":24,"news_slug":31,"published_at":32,"created_at":33,"modified_at":34,"is_published":35,"publish_type":36,"image_url":14,"view_count":37},"21c7dec1-f68e-4641-974c-ae2bce87393e","教师打分、验证器掌舵:腾讯混元 FlowBalance 给自蒸馏装上方向门控,Qwen3-8B 数学均值超 GRPO 2.12 分","腾讯混元 FlowBalance 给特权自蒸馏装上验证器符号门控:正确轨迹保留教师引导,错误轨迹反向,无偏好时关闭。Qwen3-8B 数学五榜均值 67.61,较 GRPO +2.12,训练提速 1.43 倍,并避开直接 OPSD 的响应长度坍缩。","带参考答案去重新打分自己刚生成的推理轨迹——这类「特权自蒸馏」给长推理链提供了密集的 token 级监督,恰好是纯强化学习最缺的东西。但它有个已知病灶:老师视角的打分和最终验证器的对错判定是两套信号,一条看起来头头是道、最后算错的轨迹,照样能拿到高分离子,越蒸馏越自信。腾讯混元前沿团队(HY LLM Frontier)与宾夕法尼亚大学 9 月 3 日挂出 arXiv 论文 FlowBalance,修法很直接:让验证器决定方向,自教师只负责修正幅度。\n\n## 机制:符号门控加轨迹能量\n\nFlowBalance 建立在 FlowRL 的分布视角上。每条 on-policy 轨迹先由冻结的自教师(带参考解等特权上下文)重新打分,把 token 级对数概率增益聚合成轨迹级引导分数;再与验证器给出的组相对优势合成一个轨迹能量,教师项乘以优势的符号。门控逻辑就三行:轨迹被验证为正优势,教师打分照常加权;被拒绝的轨迹,教师压力整个反向;整组无对错偏好,教师项直接关闭。之后用带剖面的轨迹平衡去拟合能量重加权后的完整回复分布,每组 rollout 只需一次对数配分估计,不再需要单独的 token 级模仿损失。论文还给出了组内对比保持、最小改动 reverse-KL 表征等性质分析。\n\n## 数字面:五榜均值与一个刺眼的对照\n\n团队在五个数学推理基准上对比 GRPO、直接 OPSD、RLSD 与 FlowRL(180 步、五种子)。Qwen3-4B 上 FlowBalance 均值 64.26,较 GRPO +1.95;Qwen3-8B 上均值 67.61,+2.12,五个分榜全部拿到最优均值,AIME24 Pass@16 达 89.33。4B 侧并非全面领先:Minerva 一项由 FlowRL 领先(51.99 对 50.51)。最刺眼的对照是直接 OPSD:4B 上均值掉 8.19 分,8B 上直接崩到 41.16,较 GRPO 掉 24.33,响应长度同步坍缩——特权模仿压短推理链的病灶被量化得明明白白。训练动力学上,FlowBalance 约 100 步达到 0.5 AIME24 验证精度,GRPO 约需 143 步(1.43 倍提速),并在 400 步内保持稳定,而 GRPO 约 180 步后明显退化。消融显示验证器系数 15、教师系数 1 为峰值配置(五榜均值 67.61)。\n\n## 多样性:不止答对,还要多路答对\n\n团队用 GPT-5.5 抽取 AIME24 完整轨迹的数学表征与工具再聚类,只统计正确轨迹的 Simpson 策略多样性:FlowBalance 0.2194,约为 GRPO(0.1017)的 2.16 倍,也高于 RLSD(0.1456)。案例里同一道四面体题,常见解法走 Cayley–Menger 行列式,FlowBalance 的轨迹则识别出 41=4²+5²、80=4²+8²、89=5²+8²,把顶点嵌进 4×5×8 长方体用混合积算体积——两条路都得到答案 104。作者也明说这是单种子、LLM 评判的可控诊断,不是种群级保证。\n\n## 冷水\n\n全部数字为团队自报,第三方复现尚未出现;实验只覆盖数学推理与 Qwen3 两个规模,代码仓库刚放出(33 个 commit),社区关注还很低。这套方法迁移到代码、agent 等没有干净验证器的领域,成本要另算。但「验证器掌舵、密集信号只做修正」这个原则,给所有想混用 RLVR 与自蒸馏的团队指了一条清晰的路:先定对错,再谈打分。\n\n参考: arXiv:2609.03241 (arxiv.org\u002Fabs\u002F2609.03241);alexhuang13.github.io\u002FFlowBalance-Blog;github.com\u002Falexhuang13\u002FFlowBalance","https:\u002F\u002Farxiv.org\u002Fabs\u002F2609.03241","7437aeb9-930c-4866-a2e9-48003c1a792b",[11,15,18,21],{"id":12,"name":13,"slug":13,"description":14,"color":14},"5e628969-6d2a-437f-998a-104e4b16cfb1","ai-progress",null,{"id":16,"name":17,"slug":17,"description":14,"color":14},"0ef8513a-0a26-42f0-b6f9-5b6dadded45c","efficiency",{"id":19,"name":20,"slug":20,"description":14,"color":14},"01598627-1ea6-4b27-a5d8-874971571a71","llm",{"id":22,"name":23,"slug":23,"description":14,"color":14},"c187600e-804c-4697-b828-1e4330e0eb10","qwen",[25],{"id":26,"lang":27,"title":28,"summary":29,"content":30},"20eada0b-0d80-48ff-b042-c5c22d3cd9a1","en","FlowBalance: verifier-gated self-distillation, +2.12 over GRPO","Tencent Hunyuan's FlowBalance sign-gates self-distillation with verifier advantage: +2.12 avg over GRPO on Qwen3-8B math, 1.43x faster training.","Rescoring your own freshly generated reasoning trajectories while holding the reference answer—privileged on-policy self-distillation supplies exactly the dense token-level supervision that pure reinforcement learning lacks. But it has a known failure mode: the teacher's score and the verifier's verdict are two different signals. A trajectory that reads coherently yet lands on a wrong answer can still collect high teacher gains, and distillation keeps reinforcing that confidence. On September 3, a team from Tencent's HY LLM Frontier group and the University of Pennsylvania posted FlowBalance to arXiv, with a blunt fix: let the verifier decide direction, and let the self-teacher only refine magnitude.\n\n## Mechanism: sign gating plus trajectory energy\n\nFlowBalance builds on FlowRL's distributional view. Each on-policy trajectory is first rescored by a frozen copy of the policy given privileged context such as a reference solution; token-level log-probability gains are clipped and averaged into a trajectory-level guidance score. That score is then combined with the verifier-derived group-relative advantage into a single trajectory energy, with the teacher term multiplied by the sign of the advantage. The gating logic is three lines: on verified positive-advantage trajectories, teacher support is retained; on rejected trajectories, teacher pressure is reversed; when the rollout group shows no outcome preference, the teacher term is disabled entirely. Profiled trajectory balance then fits the energy-reweighted distribution over complete responses, with one stopped log-partition estimate per rollout group and no separate token-level imitation loss. The paper also establishes within-group contrast preservation and a minimum-change reverse-KL characterization.\n\n## Numbers: five-benchmark averages and one glaring contrast\n\nAcross five mathematical reasoning benchmarks (step-180, five seeds), FlowBalance averages 64.26 on Qwen3-4B, +1.95 over GRPO, and 67.61 on Qwen3-8B, +2.12, taking the best mean on all five sub-benchmarks at the 8B scale with 89.33 AIME24 Pass@16. The 4B picture is not a clean sweep: FlowRL leads Minerva there, 51.99 versus 50.51. The most glaring contrast is direct OPSD: down 8.19 points on average at 4B, and a collapse to 41.16 at 8B—down 24.33 from GRPO—accompanied by shrinking response length, quantifying how privileged imitation crushes long reasoning chains. On training dynamics, FlowBalance reaches 0.5 AIME24 validation accuracy in roughly 100 steps versus about 143 for GRPO (a 1.43x speedup) and stays stable through 400 steps, while GRPO degrades sharply after around step 180. Ablations put the peak configuration at verifier coefficient 15 and teacher coefficient 1 (67.61 five-benchmark average).\n\n## Diversity: not just correct, but correct in more ways\n\nThe team used GPT-5.5 to extract the mathematical representation and tools from full AIME24 trajectories, then clustered anonymized summaries by semantic strategy and measured Simpson diversity over correct trajectories only: 0.2194 for FlowBalance, about 2.16x GRPO's 0.1017 and above RLSD's 0.1456. In one showcased tetrahedron problem, the common GRPO route runs through the Cayley–Menger determinant, while a FlowBalance trajectory recognizes 41=4²+5², 80=4²+8², and 89=5²+8², embeds the vertices in a 4×5×8 box, and computes the volume with a scalar triple product—both routes arrive at answer 104. The authors are explicit that this is a controlled, one-seed, LLM-judged diagnostic, not a population-level guarantee.\n\n## Caveats\n\nEvery number is self-reported; no third-party replication has appeared yet. Experiments cover mathematical reasoning on two Qwen3 scales only, and the code repository just landed (33 commits) with little community pickup so far. Porting the recipe to domains without clean verifiers—code, agents—carries unbudgeted cost. Still, the principle of \"verifier sets the direction, dense signals only refine\" gives every team mixing RLVR with self-distillation a clear rule: settle correctness first, then talk about scores.\n\nReferences: arXiv:2609.03241 (https:\u002F\u002Farxiv.org\u002Fabs\u002F2609.03241); project page alexhuang13.github.io\u002FFlowBalance-Blog; code github.com\u002Falexhuang13\u002FFlowBalance","flowbalance-verifier-gated-self-distillation","2026-09-08T15:08:17Z","2026-09-08T15:08:20.367776Z","2026-09-08T15:08:20.367789Z",true,"agent",101,[39,48],{"slug":40,"tag_slug":40,"title_zh":41,"title_en":42,"intro_zh":43,"intro_en":44,"id":45,"is_active":35,"created_at":46,"modified_at":47},"ai-for-science","AI for Science 2026：从 UniPert 到 GPT-Rosalind 的硬核进化","AI for Science 2026: from UniPert to GPT-Rosalind","生命科学、化学材料、物理世界模型——AI 正在从\"语言工具\"变成\"实验伙伴\"。本专题收录 AI 在三大科学方向的关键节点：UniPert 统一基因与化学扰动空间、GPT-Rosalind 端到端生命科学推理、达摩院 AI 智能体 28 小时找到 4 种超导新材料、Anthropic Claude Science 把工作台做成标准品。","From language tool to lab partner — AI is reshaping life sciences, chemistry\u002Fmaterials, and physical world models. This topic covers the key milestones: UniPert unifying genetic-chemical perturbation spaces, GPT-Rosalind's end-to-end life-sciences reasoning, DAMO's AI agent discovering 4 superconducting materials in 28 hours, and Anthropic's Claude Science workbench going mainstream.","988a4300-5fab-41c4-b5d8-63711a2dc757","2026-09-10T01:34:15.296649Z","2026-09-10T01:34:15.296663Z",{"slug":49,"tag_slug":49,"title_zh":50,"title_en":51,"intro_zh":52,"intro_en":53,"id":54,"is_active":35,"created_at":55,"modified_at":56},"h3-series","MiniMax H3 系列：从开源权重到 35 倍吞吐","MiniMax H3 Series: from open weights to 35x throughput","MiniMax H3 自 2026 年 8 月开源以来节奏密集：官方把生成、参考与编辑收回一个模型；ComfyUI 当天压进 RTX 3060；摩尔线程 3 小时完成国产 GPU 适配；fal 后训练版把吞吐拉到 35 倍；FastH3 蒸馏再砍推理成本。本专题持续追踪 H3 的发布—开源—蒸馏—部署全链路。","Since MiniMax open-sourced H3 in August 2026 the pace has been relentless: one unified omni-modal model, same-day ComfyUI support down to an RTX 3060, a 3-hour Day-0 port to Moore Threads GPUs, fal's post-trained H3 Max at 35x throughput, and FastH3 distillation cutting inference cost further. This topic tracks the full H3 chain — release, open weights, distillation, deployment.","83ef0daa-3c31-4cb3-86ed-e5ee58654d5f","2026-09-08T07:33:19.942193Z","2026-09-08T07:33:19.942209Z",{"items":58},[59,64,69,74,79,84],{"id":60,"title":61,"news_slug":62,"published_at":63},"b362eb89-32ef-46ed-b65a-dd65f6f305f2","Direct-OPD 把「RL 经验」跨模型规模可复用：字节×清华让弱模型的策略差当强模型的隐式奖励","direct-opd-rl-experience-transfer","2026-07-14T14:10:00+00:00",{"id":65,"title":66,"news_slug":67,"published_at":68},"b0c4e8d2-5662-4e3e-b489-6202eabbe97b","Dream-RSI 把历史当模拟器:162 倍杠杆重写 RSI 算力账本","dream-rsi-replay-simulator-162x","2026-09-16T06:00:00+00:00",{"id":70,"title":71,"news_slug":72,"published_at":73},"2266cea6-06f1-4932-8905-1bc3f2e5a8c0","Meta FAIR 字节蒸馏研究:End-Of-Token 渐近反超 token 蒸馏 4%,数据只需 1\u002F6","meta-fair-byte-distillation-token-ceiling-2026-09","2026-09-15T02:00:00+00:00",{"id":75,"title":76,"news_slug":77,"published_at":78},"178aa5e5-2a4f-4a87-a97c-0da16295d96f","EMNLP 2026 OCGQuant:用通道配对治 NVFP4 陪葬误差,Qwen3-1.7B 接近 FP16","ocgquant-nvfp4-outlier-companion-grouping","2026-09-10T09:15:00+00:00",{"id":80,"title":81,"news_slug":82,"published_at":83},"d941056b-c2e7-42e5-965a-a982c20b1169","Qwen3.8-Flash-Next 架构细节:Gated Residual 多分支残差 + QSA micro-block 稀疏注意力","qwen3-8-flash-next-cost-efficiency-architecture","2026-09-02T02:00:00+00:00",{"id":85,"title":86,"news_slug":87,"published_at":88},"fdbe1ee2-131c-4632-baf6-03109d7c1814","Qwen3.8-Next 架构论文:125B 参数 6B 激活,1\u002F9 训练 FLOPs 对标 397B 前辈","qwen3-8-flash-next-architecture","2026-09-01T23:15:00+00:00"]