[{"data":1,"prerenderedAt":-1},["ShallowReactive",2],{"news-slug-skillopt-microsoft-research-md-skill-trainer":3,"news-related-a988167f-3249-4854-9778-3af0b657466a":36},{"id":4,"title":5,"summary":6,"content":6,"original_url":7,"source_id":8,"tags":9,"translations":23,"news_slug":29,"published_at":30,"created_at":31,"modified_at":32,"is_published":33,"publish_type":34,"image_url":13,"view_count":35},"a988167f-3249-4854-9778-3af0b657466a","SkillOpt：把 Agent 技能 .md 文档训练成\"可学习参数\"","Microsoft Research 在 5 月发布的 SkillOpt 把 Agent 工程里一件原本靠人肉手写的事——改写「技能 .md」——变成类似深度学习训练的可控流程。核心观察是：与其继续让 Prompt 工程师手改技能文档，不如让一个外部的「优化器模型」按 SGD 的纪律去更新它。具体做法是：固定目标模型不训练，把当前技能插入 Agent 上下文采样若干轨迹；优化器把成功\u002F失败分成小批，针对失败模式提出增\u002F删\u002F改三类编辑；候选技能必须通过一个 held-out 验证门才被接受，被拒的编辑进入 rejected buffer 当作负反馈；每轮 epoch 结束再做一次「慢更新」做动量式总结。整个流程加在文本空间上，部署时零额外推理调用。\n\n论文在 6 个基准、7 个目标模型、3 种 harness 共 52 个 cell 上评估，SkillOpt 全部拿到 best-or-tied，击败或打平 no-skill、人工写、LLM 一次性写、Trace2Skill、TextGrad、GEPA、EvoSkill 等所有基线。GPT-5.5 直聊平均 +23.5 分，Codex +24.8，Claude Code +19.1；最具说服力的是跨 harness 迁移——一个在 Codex 里训练出的电子表格技能直接搬到 Claude Code，拿到 +59.7 的相对提升。\n\n传统「写 skill」是黑魔法：动一行可能把 80 分改回 40 分。SkillOpt 的工程价值在于把\"学习率（编辑预算）、验证集、动量、负反馈\"这套早已在神经网络里成熟的控件，原样搬到文本上，让 300–2000 token 的技能 .md 也能稳定训练。同一份 best_skill.md 可跨模型、跨 harness、跨相近任务复用，对闭源模型同样适用，因为它碰的是上下文而不是权重。对正在用 Codex\u002FClaude Code 搭 Agent 团队的人来说，这是个值得关注的工作流改造。","https:\u002F\u002Farxiv.org\u002Fabs\u002F2605.23904","7437aeb9-930c-4866-a2e9-48003c1a792b",[10,14,17,20],{"id":11,"name":12,"slug":12,"description":13,"color":13},"6ad31a14-c0da-42df-81fd-564281f768db","agentic-ai",null,{"id":15,"name":16,"slug":16,"description":13,"color":13},"5e628969-6d2a-437f-998a-104e4b16cfb1","ai-progress",{"id":18,"name":19,"slug":19,"description":13,"color":13},"40269b40-7942-4650-9672-ed2e6524d37a","ai-technology",{"id":21,"name":22,"slug":22,"description":13,"color":13},"b9bd9039-fcdb-41a8-b85b-fc1587def2b9","open-source",[24],{"id":25,"lang":26,"title":27,"summary":28,"content":13},"affd108d-d216-402e-a22e-8486638afb7a","en","SkillOpt: trains Agent skill .md documents into \"learnable parameters\"","arXiv 2605.23904 introduces SkillOpt, a method for converting Agent skill descriptions (typically in .md files) into learnable parameters. The result: Agents can be fine-tuned on their own skill descriptions, with the skills becoming part of the model's weights rather than just prompt context.\n\nThe technical details: SkillOpt takes a .md file describing an Agent skill (e.g., \"How to write a unit test in Python\") and converts it into a \"skill embedding\" — a continuous vector representation that can be inserted into the model. The skill embedding is trained via a \"skill distillation\" loss: the model's output with the skill embedding should match the output of a model that has been prompted with the full .md file.\n\nThe benefit: the skill becomes part of the model's \"parameter memory,\" not just \"context memory.\" This means: (1) the skill doesn't consume context window space; (2) the skill is always available, even if the .md file is missing; (3) the skill can be combined with other skills, with the model learning to use them together.\n\nThe benchmark: on a set of Agent tasks (writing unit tests, generating API documentation, fixing simple bugs), SkillOpt-trained models match the quality of prompt-based skill usage while saving 40-60% of the context window. The skill embeddings are also composable — multiple skill embeddings can be merged to handle complex tasks.\n\nThe bigger takeaway: \"skill as a parameter\" is a new abstraction for Agent systems. The traditional \"skill as a .md file\" approach has scaling limits (context window overflow, retrieval errors). SkillOpt's \"skill as a parameter\" approach is more efficient and reliable, and it paves the way for \"skill libraries\" that can be plugged into Agents at inference time. For the industry, this means Agent vendors will need to invest in \"skill embedding\" infrastructure, not just \"skill management.\"","skillopt-microsoft-research-md-skill-trainer","2026-06-21T22:01:00Z","2026-06-21T22:09:27.094375Z","2026-08-19T02:08:40.142862Z",true,"agent",81,{"items":37},[38,43,48,53,58,63],{"id":39,"title":40,"news_slug":41,"published_at":42},"777afb24-262f-45cc-961f-d5d49ad42883","AgentOPSD 用递归贝叶斯信念破解多轮 Agent 强化学习的信用分配：清华\u002F浙大\u002F美团让 GRPO 学会看哪个 turn 决定胜负","agentopsd-recursive-belief-credit-assignment","2026-08-07T02:00:00+00:00",{"id":44,"title":45,"news_slug":46,"published_at":47},"12df58ff-0771-4c0e-bf0e-00bfdc8112bb","SkillCenter：21 万可审计 Agent 技能库，SQLite 离线检索","skillcenter-sqlite-agent","2026-07-09T04:30:00+00:00",{"id":49,"title":50,"news_slug":51,"published_at":52},"deac2d55-76a6-40d2-8ef7-36aed2ad0105","Linux 7.2 把 AI 拉进内核开发:Sashiko 让补丁数量翻倍,Torvalds 接受「新常态」","linux-7-2-sashiko-ai-kernel-review","2026-08-20T12:00:00+00:00",{"id":54,"title":55,"news_slug":56,"published_at":57},"a91067a3-4fa4-4e88-a25a-18ba3bea21ea","Google 把\"加密推理\"摆上桌面：HEIR 编译器让预训练模型在密文上直接跑","google-heir-compiler-encrypted-ai-inference","2026-08-14T14:00:00+00:00",{"id":59,"title":60,"news_slug":61,"published_at":62},"22a1a718-0eb6-46e5-8ee8-825400de11d1","DeepMind WeatherNext 在 Nature 发论文：用 28 km 粗分辨率做出多一天的飓风预警,代码权重全部开源","deepmind-weathernext-cyclones-nature-open-source","2026-08-10T02:00:00+00:00",{"id":64,"title":65,"news_slug":66,"published_at":67},"9ef626d9-05dd-4e67-9069-093f90a3fd5c","Rust 主仓库正式引入 LLM 政策：把\"必须人为可读、不可代写\"写进 PR 流程","rust-lang-rust-llm-policy-adoption","2026-08-07T00:00:00+00:00"]