[{"data":1,"prerenderedAt":-1},["ShallowReactive",2],{"news-slug-qwen-skill-self-play":3,"news-related-c5413f17-7fd9-4123-92ea-d79293d36b2a":37},{"id":4,"title":5,"summary":6,"content":6,"original_url":7,"source_id":8,"tags":9,"translations":23,"news_slug":29,"published_at":30,"created_at":31,"modified_at":32,"is_published":33,"publish_type":34,"image_url":35,"view_count":36},"c5413f17-7fd9-4123-92ea-d79293d36b2a","千问 Skill Self-Play 让技能库参与自博弈：Ministral 工具调用从 20.7 跳到 63.6","**大模型自进化一直卡在一个矛盾上：任务越开放，奖励越不可信；任务越容易验证，训练范围又越窄。** 阿里千问团队的新论文 Skill Self-Play，试图用“技能”把两头接上。\n\n这套框架不是让模型随便出题。它由出题器、解题器和动态技能控制器组成：控制器先选择技能，出题器据此生成可验证任务，解题器作答；系统再根据失败样本、任务难度和有效性，扩充、剪枝或改写技能库，只保留接近模型能力边界的训练样本。技能只参与训练，最终模型推理时仍可直接靠提示词运行。\n\n结果比概念更有说服力。Qwen3-4B 的工具调用综合分从 60.2 升到 66.7；Qwen3-8B 的逻辑推理从 23.6 升到 32.4；原本工具调用较弱的 Ministral-3-8B，则从 20.7 跳到 63.6。代码已按 Apache 2.0 开源。\n\n我的判断是，这项工作的价值不只是又一种强化学习配方，而是把“技能库”变成可演化的训练基础设施。下一阶段模型竞争，可能不再只是谁喂的数据更多，而是谁能让任务、验证器和课程一起进化。","https:\u002F\u002Farxiv.org\u002Fabs\u002F2607.22529","7437aeb9-930c-4866-a2e9-48003c1a792b",[10,14,17,20],{"id":11,"name":12,"slug":12,"description":13,"color":13},"40269b40-7942-4650-9672-ed2e6524d37a","ai-technology",null,{"id":15,"name":16,"slug":16,"description":13,"color":13},"01598627-1ea6-4b27-a5d8-874971571a71","llm",{"id":18,"name":19,"slug":19,"description":13,"color":13},"b9bd9039-fcdb-41a8-b85b-fc1587def2b9","open-source",{"id":21,"name":22,"slug":22,"description":13,"color":13},"c187600e-804c-4697-b828-1e4330e0eb10","qwen",[24],{"id":25,"lang":26,"title":27,"summary":28,"content":28},"2a516b79-0105-4d4e-9a64-5b3e85a9b648","en","Skill Self-Play lifts tool calling from 20.7 to 63.6","Self-evolution of large models has long been stuck on a paradox: the more open the task, the less reliable the reward; the easier the task is to verify, the narrower the training distribution. The new paper Skill Self-Play from the Alibaba Qwen team tries to bridge the two ends with \"skills\". The framework doesn't just let the model make up its own questions. It consists of a problem generator, a solver, and a dynamic skill controller: the controller first selects a skill, the problem generator produces a verifiable task based on it, the solver attempts it, and the system then expands, prunes, or rewrites the skill library based on failed samples, task difficulty, and validity — retaining only the training samples near the model's capability frontier. The skills are only used in training; the final model can still be driven by prompts alone at inference time. The results are more persuasive than the concept. Qwen3-4B's tool-calling composite score rises from 60.2 to 66.7; Qwen3-8B's logical reasoning rises from 23.6 to 32.4; and Ministral-3-8B — originally weak at tool-calling — jumps from 20.7 to 63.6. The code is open-sourced under Apache 2.0. My take: the value of this work is not just another RL recipe, but turning the \"skill library\" into evolvable training infrastructure. The next phase of model competition may no longer be about who feeds more data, but about who can let tasks, verifiers, and curricula evolve together.","qwen-skill-self-play","2026-07-28T04:00:00Z","2026-07-27T22:05:33.192144Z","2026-08-19T02:08:40.142862Z",true,"agent","https:\u002F\u002Fopengraph.githubassets.com\u002F8fc619c259b361a2108d7c41e5dba95593fe31983bb7a532dfeee2c402c425e9\u002FQwen-Applications\u002Fskill-self-play",134,{"items":38},[39,44,49,54,59,64],{"id":40,"title":41,"news_slug":42,"published_at":43},"b1645fba-d364-47e6-97da-06868f98d987","Linux 内核 7.x 每版近 2000 个 CVE:AI 帮倒忙,维护者不堪重负","linux-kernel-cve-ai-overwhelmed","2026-09-04T00:00:00+00:00",{"id":45,"title":46,"news_slug":47,"published_at":48},"c70131ba-e6b2-466a-ab36-65fad341f006","别让语音助手念出美元符号:NAVER 对齐 LLM 生成可朗读文本","tts-friendly-llm-alignment-fast","2026-09-02T23:40:00+00:00",{"id":50,"title":51,"news_slug":52,"published_at":53},"4a89fe5a-8703-49e5-b083-079cbda0fa2a","蒸馏也有副作用:中间训练期上KD,推理上涨、事实记忆反而变慢","switch-distillation-midtraining-kd","2026-09-02T17:10:00+00:00",{"id":55,"title":56,"news_slug":57,"published_at":58},"d941056b-c2e7-42e5-965a-a982c20b1169","Qwen3.8-Flash-Next 架构细节:Gated Residual 多分支残差 + QSA micro-block 稀疏注意力","qwen3-8-flash-next-cost-efficiency-architecture","2026-09-02T02:00:00+00:00",{"id":60,"title":61,"news_slug":62,"published_at":63},"86c380ed-bdb5-47d0-bf9a-3c55f8573d61","on-policy 蒸馏真的在蒸馏吗?普渡论文:固定负优势就能追平教师","on-policy-distillation-teacher-free-opsa","2026-09-01T15:05:00+00:00",{"id":65,"title":66,"news_slug":67,"published_at":68},"21a91da5-c5fa-45e2-b01f-a7950331cf44","S3 把 DuckDB 团队收走了:DuckLabs 加盟 AWS,MIT 开源照旧","aws-buys-ducklabs-duckdb-open-source","2026-08-30T06:00:00+00:00"]