[{"data":1,"prerenderedAt":-1},["ShallowReactive",2],{"news-slug-bcit-conditional-experience-transfer-post-training":3,"topics-all":38,"news-related-7623f190-7071-4811-a6f1-32462a99b8d3":48},{"id":4,"title":5,"summary":6,"content":7,"original_url":8,"source_id":9,"tags":10,"translations":24,"news_slug":31,"published_at":32,"created_at":33,"modified_at":34,"is_published":35,"publish_type":36,"image_url":14,"view_count":37},"7623f190-7071-4811-a6f1-32462a99b8d3","经验会过期:阿里云论文让自主后训练的有害授权率从 62.5% 降到 25%","模型自主后训练会复用历史成功经验,但父模型变了,旧经验可能失效甚至有害。阿里云团队的 BCIT 给每次复用加「拒绝-验证-训练」闸门,在 Qwen3-4B 上把有害候选授权率从 62.5% 压到 25%,等算力下最终成绩反超最强基线 2.63 分。","模型上线之后,后训练不会停:新领域、新工具、新需求,都要往同一个基座上叠更新。当这一过程交给自主系统去跑——提方案、训候选、按评测反馈挑下一个——一个新问题浮出水面:过去某次更新确实有效,这份「成功经验」在基座又被改过几轮之后,还有多少能直接复用?阿里云团队的一篇新论文给出了一个反直觉的答案:先问该不该复用,再谈怎么复用。\n\n## 问题:经验会过期\n\narXiv 2608.26730 把这个场景形式化为「条件经验迁移」。核心观察:一次更新的效果取决于它的父模型、数据和训练阶段——三者任何一个变了,旧结论都可能失效。把历史成功当成无条件的通行证,轻则浪费算力,重则让被提升的子模型带伤上岗,污染后续整条训练轨迹。\n\n论文的实验直接量化了这种「异质性」:同一个候选更新,在不同上下文里的效果差异巨大。例如金融领域的 precision replay 更新,在其源任务上小幅提升 0.35,却在迁移目标上带来 +13.3 的跨度;而 SQL 的 rationale-first SFT 换个上下文反而倒扣 2.25。经验不是资产,是易耗品。\n\n## 方法:先过闸门,再给算力\n\n团队提出 Boundary-Calibrated Intervention Transfer(BCIT),一个插在「候选生成」和「全量训练」之间的授权层。它的动作只有三种:拒绝(Reject)、有界验证(Validate)、训练(Train)。\n\n规则设计得很克制:每条历史经验都绑定它当初生效的源上下文;复用前核对预设适用条件;命名列名的硬冲突直接否决——比如一个函数调用更新撞上不兼容的输出协议,哪怕正分再高也不补偿;证据不足时,只在当前父模型上跑一次有界的小规模试训取证。新提案没有历史战绩,一律不许绕过验证。乘法门控(源强度 × 当前兼容度)保证源证据再强,也补不了上下文的错位。\n\n## 数字:有害授权砍掉六成\n\n在 Qwen3-4B 上跨金融推理、text-to-SQL、函数调用三个方向的连续适配实验里,BCIT 的成绩单有三组数字值得记住。\n\n其一,24 个候选的事后盲审(Audit-24,含 10 个有益、8 个有害、6 个中性)中,BCIT 只放行了 8 个有害候选中的 2 个,有害授权率 25.0%;对照的 Flat-Additive 打分基线放行了 5 个,授权率 62.5%——砍掉六成。同时它保住了 10 个有益候选中的 9 个(90.0%),不是靠一刀切拒掉了事。\n\n其二,在候选流、评测、36 GPU 时预算全部对齐的六轮配对实验中,BCIT 最终跨任务均值 47.0,比 Flat-Additive 高 2.63 分(95% CI [2.10, 3.16],六轮全胜),也压过 Validate-All(1.50 分)和 Additive+Veto(0.90 分)。\n\n其三,晋升率:BCIT 训完的 62 个候选里 46 个被采纳(74.2%),Flat-Additive 只有 44.4%——多出来的那一半,正是被错误授权的有害更新挤占的名额。\n\n训练栈本身很朴素:Qwen3-4B,LlamaFactory + PyTorch,DeepSpeed ZeRO-2,bfloat16,4096 上下文——没有堆任何特殊基础设施,闸门收益不依赖工程红利。\n\n## 所以呢\n\n自主后训练的叙事里,大家默认难点在「怎么更快地试更多方案」。这篇论文指出另一个同样重要的轴:在算力花出去之前判断「这次不该复用」,本身就是一个独立的技术问题。对做持续迭代模型平台的团队,这个「经验授权层」值得进架构图——毕竟拒绝一次有害更新省下的,不只是算力,还有被污染的训练轨迹。\n\n参考:arXiv:2608.26730(https:\u002F\u002Farxiv.org\u002Fabs\u002F2608.26730)","https:\u002F\u002Farxiv.org\u002Fabs\u002F2608.26730","7437aeb9-930c-4866-a2e9-48003c1a792b",[11,15,18,21],{"id":12,"name":13,"slug":13,"description":14,"color":14},"40269b40-7942-4650-9672-ed2e6524d37a","ai-technology",null,{"id":16,"name":17,"slug":17,"description":14,"color":14},"0ef8513a-0a26-42f0-b6f9-5b6dadded45c","efficiency",{"id":19,"name":20,"slug":20,"description":14,"color":14},"01598627-1ea6-4b27-a5d8-874971571a71","llm",{"id":22,"name":23,"slug":23,"description":14,"color":14},"b9bd9039-fcdb-41a8-b85b-fc1587def2b9","open-source",[25],{"id":26,"lang":27,"title":28,"summary":29,"content":30},"1ee9e657-6b87-4bae-b698-49107ab4670f","en","Alibaba Cloud's BCIT Cuts Harmful Reuse in Autonomous Post-Training","Alibaba Cloud's BCIT gates reuse of past experience in autonomous post-training, cutting harmful authorization from 62.5% to 25% on Qwen3-4B.","When models ship, post-training does not stop: new domains, new tools, and new requirements all pile updates onto the same base. Hand that loop to an autonomous system — propose updates, train candidates, pick the next one by evaluation feedback — and a new question emerges: an update that demonstrably worked in the past may no longer be reusable after the parent model has been changed several times. A new paper from an Alibaba Cloud team gives a counterintuitive answer: ask whether to reuse at all, before asking how.\n\n## The Problem: Experience Expires\n\narXiv 2608.26730 formalizes this setting as \"conditional experience transfer.\" The core observation: an update's effect depends on its parent model, data, and training stage — change any of the three, and the old conclusion can break. Treating past success as context-free permission wastes compute at best; at worst the promoted child ships with damage and contaminates the rest of the training trajectory.\n\nThe paper quantifies this heterogeneity directly. The same candidate update swings wildly across contexts: a finance-domain precision replay update gains a modest +0.35 on its source task yet delivers a +13.3 swing on a transfer target, while the SQL rationale-first SFT loses 2.25 when the context shifts. Experience is not an asset — it is a perishable.\n\n## The Method: A Gate Before Compute\n\nThe team proposes Boundary-Calibrated Intervention Transfer (BCIT), an authorization layer that sits between candidate generation and full weight-changing training. It takes exactly three actions: Reject, Validate, or Train.\n\nThe rules are deliberately restrained. Every piece of historical evidence is bound to the source context where its effect was observed; before reuse, prespecified applicability conditions are checked; candidates with named hard conflicts are vetoed outright — a function-calling update hitting an incompatible output protocol is non-compensable, no matter how positive the score. When evidence is insufficient, BCIT runs one bounded trial on the current parent to gather fresh evidence. New proposals carry no track record and can never bypass validation. A multiplicative gate (source strength × current compatibility) ensures that strong source evidence cannot compensate for a context mismatch.\n\n## The Numbers: Harmful Authorization Cut by Sixty Percent\n\nOn Qwen3-4B, adapted sequentially across finance reasoning, text-to-SQL, and function calling, three groups of results stand out.\n\nFirst, in the outcome-blind audit of 24 candidates (10 beneficial, 8 harmful, 6 neutral), BCIT authorized only 2 of the 8 harmful candidates — a 25.0% harmful-authorization rate versus 62.5% for the Flat-Additive scoring baseline. It still kept 9 of the 10 beneficial candidates (90.0%), so the win is not a trivial reject-everything policy.\n\nSecond, across six paired episodes with identical candidate streams, evaluations, and a 36-GPU-hour cap, BCIT's final cross-task mean of 47.0 beats Flat-Additive by 2.63 points (95% CI [2.10, 3.16], positive in all six pairs), and also exceeds Validate-All (by 1.50) and Additive+Veto (by 0.90).\n\nThird, promotion yield: 46 of BCIT's 62 fully trained candidates were adopted (74.2%) versus 44.4% for Flat-Additive — the extra half is precisely the slots that wrongfully authorized harmful updates would have occupied.\n\nThe training stack itself is plain: Qwen3-4B, LlamaFactory + PyTorch, DeepSpeed ZeRO-2, bfloat16, 4,096-token context — no exotic infrastructure, which suggests the gate's gains do not ride on engineering privileges.\n\n## So What\n\nThe autonomous post-training narrative assumes the hard part is trying more candidates faster. This paper points at an equally important axis: deciding \"this should not be reused\" before the compute is spent is an independent technical problem. For teams running continuously-updated model platforms, this experience-authorization layer deserves a spot on the architecture diagram — what a rejected harmful update saves is not just compute, but the training trajectory it would have polluted.\n\nReference: arXiv:2608.26730 (https:\u002F\u002Farxiv.org\u002Fabs\u002F2608.26730)","bcit-conditional-experience-transfer-post-training","2026-09-05T17:11:11Z","2026-09-05T17:11:23.981262Z","2026-09-05T17:11:23.981271Z",true,"agent",52,[39],{"slug":40,"tag_slug":40,"title_zh":41,"title_en":42,"intro_zh":43,"intro_en":44,"id":45,"is_active":35,"created_at":46,"modified_at":47},"h3-series","MiniMax H3 系列：从开源权重到 35 倍吞吐","MiniMax H3 Series: from open weights to 35x throughput","MiniMax H3 自 2026 年 8 月开源以来节奏密集：官方把生成、参考与编辑收回一个模型；ComfyUI 当天压进 RTX 3060；摩尔线程 3 小时完成国产 GPU 适配；fal 后训练版把吞吐拉到 35 倍；FastH3 蒸馏再砍推理成本。本专题持续追踪 H3 的发布—开源—蒸馏—部署全链路。","Since MiniMax open-sourced H3 in August 2026 the pace has been relentless: one unified omni-modal model, same-day ComfyUI support down to an RTX 3060, a 3-hour Day-0 port to Moore Threads GPUs, fal's post-trained H3 Max at 35x throughput, and FastH3 distillation cutting inference cost further. This topic tracks the full H3 chain — release, open weights, distillation, deployment.","83ef0daa-3c31-4cb3-86ed-e5ee58654d5f","2026-09-08T07:33:19.942193Z","2026-09-08T07:33:19.942209Z",{"items":49},[50,55,60,65,70,75],{"id":51,"title":52,"news_slug":53,"published_at":54},"58ed753e-ad6d-4aac-95f4-36bf217e169c","把 10 万条人类视频变成机器人教材:RoboTok 检索 mAP 提升约 50 倍,hard 任务 79.3% 对 19.5%","robotok-retrieval-benchmark-reread","2026-09-06T21:11:25+00:00",{"id":56,"title":57,"news_slug":58,"published_at":59},"005557c5-8a3c-4d34-89bc-35d5351c4570","蒸馏只需要一条训练样本?清华实测:单条query覆盖71.5%训练状态,16条追平17k全量","one-shot-opd-single-query-distillation","2026-09-05T21:07:11+00:00",{"id":61,"title":62,"news_slug":63,"published_at":64},"4a89fe5a-8703-49e5-b083-079cbda0fa2a","蒸馏也有副作用:中间训练期上KD,推理上涨、事实记忆反而变慢","switch-distillation-midtraining-kd","2026-09-02T17:10:00+00:00",{"id":66,"title":67,"news_slug":68,"published_at":69},"3c9e4d7f-6f2a-4fea-8d22-c351b8fd7a4a","IBM Granite 4.1：Dense架构回归，8B参数挑战32B MoE性能","ibm-granite-4-1-dense-8b-moe-32b-grc","2026-04-29T19:10:00+00:00",{"id":71,"title":72,"news_slug":73,"published_at":74},"e84fe968-5d86-4247-baad-5da23efef860","UltraData-RL-2609 开源:85,995 条可验证奖励任务,拆解 MiniCPM5-2B 的 RL 燃料","ultradata-rl-2609-verifiable-rl-dataset","2026-09-07T23:07:45+00:00",{"id":76,"title":77,"news_slug":78,"published_at":79},"0c29e1ad-914a-4b79-a153-445c087acb03","被 LLM 抛弃的 dropout 翻身:Cerebras 称调好可省 25% 训练 FLOPs","dont-drop-dropout-layer-sparsity","2026-09-07T21:06:35+00:00"]