[{"data":1,"prerenderedAt":-1},["ShallowReactive",2],{"news-slug-qwq-32b-rl-rival-deepseek-r1-agent":3,"news-related-7b26392a-4bb3-448f-881a-335ad8aee605":36},{"id":4,"title":5,"summary":6,"content":6,"original_url":7,"source_id":8,"tags":9,"translations":23,"news_slug":29,"published_at":30,"created_at":31,"modified_at":32,"is_published":33,"publish_type":34,"image_url":13,"view_count":35},"7b26392a-4bb3-448f-881a-335ad8aee605","QwQ-32B：强化学习驱动的开源推理模型，以320亿参数比肩6710亿的DeepSeek-R1","今年1月，DeepSeek R1凭借6710亿参数和强化学习的突破在推理能力上惊艳全场。如今，阿里Qwen团队发布了QwQ-32B——一个仅有320亿参数的推理模型，却在各项基准测试中实现了与DeepSeek R1相当的性能表现。\n\n这项突破的关键在于Scaling RL策略。与R1一样，QwQ-32B采用冷启动检查点，通过基于结果的奖励在数学和编码领域进行训练，利用准确度验证器和代码执行服务器评估解决方案质量。随着训练推进，数学和编码能力持续提升，随后又增加了通用能力强化学习阶段，进一步扩展模型的泛化能力。\n\n第二阶段的RL仅需较少步数就能增强指令遵循和人类偏好对齐等通用能力，同时不显著牺牲数学和编码性能。\n\n更值得注意的是，QwQ-32B将Agent能力融入推理模型，使其能在推理过程中调用工具并根据环境反馈进行自适应调整。这是迈向Agent化推理的重要一步——推理不再只是「思考」，而是能真正「行动」。\n\n作为开源模型（Apache 2.0许可），QwQ-32B在32B参数规模下实现了与6710亿参数DeepSeek R1相当的性能，展示了强化学习与模型规模之间更优的权衡效率。它证明了推理能力的提升不一定需要成比例的参数增长，为开源社区小规模模型的推理能力提升指明了方向。\n\n但挑战同样存在：强化学习训练过程的不稳定性、对超参数的高度敏感性，以及可复现性问题，都是这类方法走向工程化部署需要解决的难题。","https:\u002F\u002Fqwenlm.github.io\u002Fblog\u002Fqwq-32b\u002F","c36a21ac-2a77-421b-9519-1e150695732a",[10,14,17,20],{"id":11,"name":12,"slug":12,"description":13,"color":13},"5e628969-6d2a-437f-998a-104e4b16cfb1","ai-progress",null,{"id":15,"name":16,"slug":16,"description":13,"color":13},"0a93ec8e-ea39-4693-81de-563ca8c173f7","inference",{"id":18,"name":19,"slug":19,"description":13,"color":13},"01598627-1ea6-4b27-a5d8-874971571a71","llm",{"id":21,"name":22,"slug":22,"description":13,"color":13},"b9bd9039-fcdb-41a8-b85b-fc1587def2b9","open-source",[24],{"id":25,"lang":26,"title":27,"summary":28,"content":13},"c95689fa-e52e-4784-8d9e-66a3e14df1d2","en","QwQ-32B: RL-trained open reasoning, matches 671B DeepSeek-R1","Back in January, DeepSeek R1 stunned the world with its 671-billion-parameter reasoning breakthrough powered by reinforcement learning. Now, Alibaba's Qwen team has released QwQ-32B — a reasoning model with only 32 billion parameters that nonetheless matches DeepSeek R1's performance across benchmarks.\n\nThe key to this breakthrough is a Scaling RL strategy. Like R1, QwQ-32B uses a cold-start checkpoint and trains in math and coding domains with outcome-based rewards, leveraging accuracy verifiers and code-execution servers to evaluate solution quality. As training progresses, math and coding abilities keep improving; later a general-capability RL stage is added to further extend the model's generalization.\n\nThe second-stage RL takes only a few steps to enhance general abilities like instruction-following and human-preference alignment, without significantly sacrificing math and coding performance.\n\nMore notably, QwQ-32B weaves agent capability into the reasoning model, letting it call tools during reasoning and adapt based on environmental feedback. This is an important step toward agentic reasoning — reasoning is no longer just \"thinking,\" it can truly \"act.\"\n\nAs an open-source model (Apache 2.0 licensed), QwQ-32B achieves performance on par with 671-billion-parameter DeepSeek R1 at 32B scale, demonstrating a more efficient trade-off between reinforcement learning and model scale. It proves that reasoning-capability gains don't necessarily require proportional parameter growth, pointing the way for open-source community efforts to boost reasoning in smaller models.\n\nBut challenges remain: instability in RL training, high hyperparameter sensitivity, and reproducibility issues are all hurdles on the path to engineered deployment.","qwq-32b-rl-rival-deepseek-r1-agent","2026-05-23T14:08:00Z","2026-05-23T22:07:57.889109Z","2026-08-19T02:08:40.142862Z",true,"agent",108,{"items":37},[38,43,48,53,58,63],{"id":39,"title":40,"news_slug":41,"published_at":42},"1311adb6-dc19-41a7-a188-6760d9e53672","HF Summer 2026 报告:13 个下载量 Top 25 模型是 2022 年的老面孔","hugging-face-summer-2026-attention-adoption","2026-08-24T08:00:00+00:00",{"id":44,"title":45,"news_slug":46,"published_at":47},"cac485ab-a429-4ebc-88c6-a1f924f978ff","AWS 开源 KeysAndValues:微调时就让模型学会“遗忘”,单张 A100 撑住 128K","aws-keysvalues-sparse-attention-finetuning","2026-08-26T05:20:00+00:00",{"id":49,"title":50,"news_slug":51,"published_at":52},"0d8fdf45-4585-47c0-9e78-3652e318b156","Apple Intelligence 中国版落地:通义千问接管语言 AI,百度负责视觉搜索","apple-intelligence-china-qwen-baidu-2026","2026-08-25T12:00:00+00:00",{"id":54,"title":55,"news_slug":56,"published_at":57},"c94bdf86-5de9-49fe-8c98-0f5c47611bfe","SGLang v0.5.18 发布:大模型冷启动提速 2.38 倍,710 个 PR 都改了什么","sglang-v0-5-18-cold-start-2-38x","2026-08-24T23:15:00+00:00",{"id":59,"title":60,"news_slug":61,"published_at":62},"988bbfb8-672a-4c6c-98f7-3a170b6bd8b3","Macaw 把 LFM2.5 装进 1.5GB:4-bit 端侧 LLM 跑 Mac 控制工具链","macaw-lfm25-15gb-edge-mac-agent","2026-08-24T06:00:00+00:00",{"id":64,"title":65,"news_slug":66,"published_at":67},"1844afb1-3a1c-4acd-9e4c-f5e2792a2018","下载免费不等于商用免费：HF Summer 2026 隐藏的开源前沿许可证分水岭","frontier-license-shift-hf-summer-2026","2026-08-23T12:30:00+00:00"]