[{"data":1,"prerenderedAt":-1},["ShallowReactive",2],{"news-slug-pistis-idrl-interleaved-distillation-rl":3,"topics-all":38,"news-related-4f5c4072-660c-47e7-9d9a-9782eb5591bf":57},{"id":4,"title":5,"summary":6,"content":7,"original_url":8,"source_id":9,"tags":10,"translations":24,"news_slug":31,"published_at":32,"created_at":33,"modified_at":34,"is_published":35,"publish_type":36,"image_url":14,"view_count":37},"4f5c4072-660c-47e7-9d9a-9782eb5591bf","Pistis 报告:IDRL 让蒸馏和 RL 交替上岗","字节跳动团队发布 Pistis 技术报告:在 Qwen3.6 和 Qwen3.5 底座上训练 27B 与 9B 两个尺寸的多模态模型,提出 IDRL 后训练范式,让在线蒸馏与强化学习在同一循环里交替进行、互相补充,缓解纯强化学习常见的策略熵坍缩;多模态搜索收益最大,一项基准较底座提升 8.8 分。","多模态后训练里,蒸馏和强化学习通常被当成前后相接的两个阶段:先蒸馏打底,再 RL 冲分。字节跳动团队的 Pistis 技术报告给出了不同答案——IDRL(Interleaved Distillation and Reinforcement Learning)把两件事放进同一个训练循环,交替进行,而不是串行接力或静态加权求和。\n\n## 先看模型本身\n\nPistis 家族包含 27B 和 9B 两档多模态模型,分别构建在 Qwen3.6-27B 和 Qwen3.5-9B 之上。每档又拆成两个变体:Pistis-Thinking 主攻深度多模态推理,Pistis-Agentic 额外吃进智能体轨迹数据,面向长程规划、迭代推理和工具调用。SFT 阶段用了约 320 万条多模态 QA 对,其中智能体轨迹按工具集成推理约 40%、搜索约 20%、通用智能体约 40% 配比。\n\n成绩上,24 个非 grounding 基准里 Pistis-27B-Thinking 与 Qwen3.8-27B 打平(82.3 对 82.4),但 grounding 均分拿到 80.5,高出 Qwen3.8-27B 8.6 分。Agentic 版本的多模态搜索尤其突出:相对 Qwen3.6-27B,BrowseComp-VL、MMSearch、VDR-testmini、LiveVQA 分别提升 8.8、3.3、3.2、9.7 分。\n\n## IDRL:不是加权,是轮班\n\n核心思路一句话:RL 让策略向高奖励输出收敛,熵下降;在线蒸馏把学生拉回教师分布,保住熵。两个目标在同一步里直接相加,梯度会打架——论文给出了推导:当学生对某个 token 已经比教师更自信时,两项贡献方向相反,合并更新甚至可能让奖励目标不增反降。IDRL 的做法是按周期轮换,比如每周期 OPD 和 RL 各 5 步,每一步只优化一个目标;等熵进入平台期,再切换到纯 RL 收尾。9B 版本用 IDRL,27B 版本用纯 RL,并且顺手当 9B 的蒸馏教师——27B-Thinking 是 9B-Thinking 的老师,27B-Agentic 教 9B-Agentic。\n\n长程智能体任务还有个专属设计:PAS(步级正优势抑制)。轨迹级奖励只看最终成败,容易让「侥幸成功」的废动作拿到正分。Pistis 把被拒绝的、重复的、答完题还在调工具的无效中间步骤的正优势直接清零,失败轨迹里照样保留惩罚,信用分配干净了不少。\n\n## PAH:不改参数,优化 harness\n\n报告还有个系统层彩蛋:PAH(Pistis-Auto-Harnessing)。模型和工具接口冻结,优化的对象是外层 harness——一个 Optimization Agent 在开发集上跑五阶段闭环:归因失败、提出一个可证伪的修改、实现、金丝雀验证、全量评估,开发指标不涨就回滚。最终产物包含 Candidate Ledger(结构化候选账本)、按需加载的 Search Skills、证据驱动的检查点和预算感知收敛,在 VDR-testmini 上验证有效,而且没有增加交互预算。\n\n## 所以呢\n\nIDRL 的价值不在「蒸馏 + RL」这个组合本身,而在「怎么组合」:交替而非加权,本质是把两个方向相反的梯度目标在时间轴上错开,让熵成为可再生的探索资源。对做后训练的团队,这是一份可复用的调度方案;对围观者,值得记住的信号是:后训练的竞争已经从「用什么数据」前移到「怎么编排训练信号」,harness 本身正在变成优化目标。([arXiv:2609.28554](https:\u002F\u002Farxiv.org\u002Fabs\u002F2609.28554))","https:\u002F\u002Farxiv.org\u002Fabs\u002F2609.28554","7437aeb9-930c-4866-a2e9-48003c1a792b",[11,15,18,21],{"id":12,"name":13,"slug":13,"description":14,"color":14},"6ad31a14-c0da-42df-81fd-564281f768db","agentic-ai",null,{"id":16,"name":17,"slug":17,"description":14,"color":14},"120fa59a-ff6f-4537-9bf5-f818df636a0e","benchmark",{"id":19,"name":20,"slug":20,"description":14,"color":14},"01598627-1ea6-4b27-a5d8-874971571a71","llm",{"id":22,"name":23,"slug":23,"description":14,"color":14},"499f4b56-819d-49a3-9609-33e775143b86","multimodal",[25],{"id":26,"lang":27,"title":28,"summary":29,"content":30},"e1834a3e-d270-4187-9d62-8a7dd77f6f12","en","Pistis report: IDRL alternates distillation and RL","ByteDance Pistis report: 27B\u002F9B multimodal models on Qwen3.6\u002F3.5; IDRL interleaves distillation and RL to curb entropy collapse; BrowseComp-VL +8.8.","Multimodal post-training usually treats distillation and reinforcement learning as two sequential stages: distill first, then RL. The Pistis technical report from a ByteDance team proposes a different answer—IDRL (Interleaved Distillation and Reinforcement Learning) puts both objectives into a single training loop and alternates between them, instead of chaining the stages or merging them into one static weighted loss.\n\n## The models first\n\nThe Pistis family covers two multimodal scales, 27B and 9B, built on Qwen3.6-27B and Qwen3.5-9B respectively. Each scale ships two variants: Pistis-Thinking targets deep multimodal reasoning, while Pistis-Agentic additionally consumes agentic trajectory data for long-horizon planning, iterative reasoning, and tool use. The SFT stage uses roughly 3.2M multimodal QA pairs, with agentic trajectories split into about 40% tool-integrated reasoning, 20% search, and 40% general agent trajectories.\n\nOn 24 non-grounding benchmarks, Pistis-27B-Thinking is comparable to Qwen3.8-27B (82.3 vs 82.4), while achieving the highest grounding average of 80.5, exceeding Qwen3.8-27B by 8.6 points. The Agentic variants are particularly strong in multimodal search: relative to Qwen3.6-27B, Pistis-27B-Agentic improves BrowseComp-VL, MMSearch, VDR-testmini, and LiveVQA by 8.8, 3.3, 3.2, and 9.7 points respectively.\n\n## IDRL: alternation, not weighting\n\nThe core idea in one sentence: RL sharpens the policy toward high-reward outputs and reduces entropy, while on-policy distillation pulls the student back toward the teacher's distribution and preserves entropy. Summing the two objectives in the same step makes their gradients fight—the report derives that when the student is already more confident than the teacher on a token, the two contributions point in opposite directions, and the summed update can even lower the reward objective. IDRL instead alternates by cycle, for example 5 OPD steps and 5 RL steps per cycle, optimizing exactly one objective per step; once entropy plateaus, it switches to pure RL. The 9B variants use IDRL while the 27B variants are trained with pure RL—and double as frozen OPD teachers: 27B-Thinking teaches 9B-Thinking, and 27B-Agentic teaches 9B-Agentic.\n\nFor long-horizon agentic tasks there is a dedicated design: PAS (step-level positive-advantage suppression). Trajectory-level rewards only see final success, which lets lucky-but-useless intermediate actions earn positive credit. Pistis zeroes the positive advantages of rejected, repeated, or post-answer tool-call steps, while keeping penalties in failed trajectories—cleaner credit assignment.\n\n## PAH: optimize the harness, not the parameters\n\nThe report also has a system-level bonus: PAH (Pistis-Auto-Harnessing). The model and tool interface stay frozen; the optimization target is the surrounding inference harness. An Optimization Agent runs a five-stage closed loop on a development set: attribute failures, propose one falsifiable change, implement it, gate it with a small canary run, then evaluate on the full development set—rolling back whenever the development metric does not improve. The resulting harness ships a structured Candidate Ledger, conditionally loaded Search Skills, evidence-driven checkpoints, and budget-aware convergence. It is validated on VDR-testmini under the same interaction budget, improving performance without extra model updates or added interaction budget.\n\n## So what\n\nThe value of IDRL is not the \"distillation plus RL\" combination itself, but how the two are combined: alternating rather than weighting is essentially about separating two opposing gradient objectives in time, making entropy a renewable exploration resource. For post-training teams, this is a reusable scheduling recipe; for observers, the signal worth remembering is that post-training competition has shifted from \"what data to use\" toward \"how to orchestrate training signals\"—and the harness itself is becoming an optimization target. ([arXiv:2609.28554](https:\u002F\u002Farxiv.org\u002Fabs\u002F2609.28554))","pistis-idrl-interleaved-distillation-rl","2026-09-25T19:11:33Z","2026-09-25T19:11:33.922305Z","2026-09-25T19:11:33.922313Z",true,"agent",296,[39,48],{"slug":40,"tag_slug":40,"title_zh":41,"title_en":42,"intro_zh":43,"intro_en":44,"id":45,"is_active":35,"created_at":46,"modified_at":47},"ai-for-science","AI for Science 2026：从 UniPert 到 GPT-Rosalind 的硬核进化","AI for Science 2026: from UniPert to GPT-Rosalind","生命科学、化学材料、物理世界模型——AI 正在从\"语言工具\"变成\"实验伙伴\"。本专题收录 AI 在三大科学方向的关键节点：UniPert 统一基因与化学扰动空间、GPT-Rosalind 端到端生命科学推理、达摩院 AI 智能体 28 小时找到 4 种超导新材料、Anthropic Claude Science 把工作台做成标准品。","From language tool to lab partner — AI is reshaping life sciences, chemistry\u002Fmaterials, and physical world models. This topic covers the key milestones: UniPert unifying genetic-chemical perturbation spaces, GPT-Rosalind's end-to-end life-sciences reasoning, DAMO's AI agent discovering 4 superconducting materials in 28 hours, and Anthropic's Claude Science workbench going mainstream.","988a4300-5fab-41c4-b5d8-63711a2dc757","2026-09-10T01:34:15.296649Z","2026-09-10T01:34:15.296663Z",{"slug":49,"tag_slug":49,"title_zh":50,"title_en":51,"intro_zh":52,"intro_en":53,"id":54,"is_active":35,"created_at":55,"modified_at":56},"h3-series","MiniMax H3 系列：从开源权重到 35 倍吞吐","MiniMax H3 Series: from open weights to 35x throughput","MiniMax H3 自 2026 年 8 月开源以来节奏密集：官方把生成、参考与编辑收回一个模型；ComfyUI 当天压进 RTX 3060；摩尔线程 3 小时完成国产 GPU 适配；fal 后训练版把吞吐拉到 35 倍；FastH3 蒸馏再砍推理成本。本专题持续追踪 H3 的发布—开源—蒸馏—部署全链路。","Since MiniMax open-sourced H3 in August 2026 the pace has been relentless: one unified omni-modal model, same-day ComfyUI support down to an RTX 3060, a 3-hour Day-0 port to Moore Threads GPUs, fal's post-trained H3 Max at 35x throughput, and FastH3 distillation cutting inference cost further. This topic tracks the full H3 chain — release, open weights, distillation, deployment.","83ef0daa-3c31-4cb3-86ed-e5ee58654d5f","2026-09-08T07:33:19.942193Z","2026-09-08T07:33:19.942209Z",{"items":58},[59,64,69,74,79,84],{"id":60,"title":61,"news_slug":62,"published_at":63},"e98cf9a2-8348-4142-9a56-c11166774798","SpeakerMem-R1:多方对话记忆,分清谁说了什么","speakermem-r1-multi-party-memory","2026-09-24T19:05:00+00:00",{"id":65,"title":66,"news_slug":67,"published_at":68},"aa9b279e-07c6-4d0c-86f0-403c4af321fa","蚂蚁 RULER:六维评分表给 SVG 生成重写奖励信号","ruler-rubric-rewards-svg-generation","2026-09-23T23:06:24+00:00",{"id":70,"title":71,"news_slug":72,"published_at":73},"832948d8-aa11-4589-b406-07ffc09eccd0","微软Taste-Bench:502个决策岔口,最强模型也只答对59.7%","taste-bench-agent-decision-forks","2026-09-23T13:10:11+00:00",{"id":75,"title":76,"news_slug":77,"published_at":78},"5d3c50e8-5087-43e9-a8f1-c64973f712c1","Qwen 拆掉 ASR 管道:音视频原生对话靠合成数据练成","qwen-omnivchat-native-audio-visual-dialogue","2026-09-21T15:14:15+00:00",{"id":80,"title":81,"news_slug":82,"published_at":83},"3bcb0e1d-99ea-4fae-9bdd-b6b625aabf10","代码 agent 8 成都在骗你:12 模型实测揭晓","overclaimbench-llm-agents","2026-09-21T07:00:00+00:00",{"id":85,"title":86,"news_slug":87,"published_at":88},"ee7c1b35-e8cc-41e3-8b8f-27a512a9f639","TempCloze 视频「完形填空」:31 款 Video-LLM 横评,开源模型时间对齐平均 26.54% 逼近乱猜","tempcloze-video-llm-temporal-alignment","2026-09-13T23:06:41+00:00"]