[{"data":1,"prerenderedAt":-1},["ShallowReactive",2],{"news-slug-ibm-granite-4-2-native-reasoning-agents":3,"topics-all":38,"news-related-71dde565-b87a-4470-83a5-0ad7a4ee1787":57},{"id":4,"title":5,"summary":6,"content":7,"original_url":8,"source_id":9,"tags":10,"translations":24,"news_slug":31,"published_at":32,"created_at":33,"modified_at":34,"is_published":35,"publish_type":36,"image_url":14,"view_count":37},"71dde565-b87a-4470-83a5-0ad7a4ee1787","IBM Granite 4.2 开源:原生思维链做成开关,30B 拿下 SWE Bench Pro","IBM 发布 Granite 4.2 开源语言模型,3B\u002F8B\u002F30B 三尺寸 dense 架构,引入原生思维链推理,支持三档思考力度切换;30B 版在 SWE Bench Pro 拿到同级最高分,Apache 2.0 许可可商用。","当大模型的竞争焦点从\"聊天\"转向\"干活\",IBM 给出了自己的答案:发布 Granite 4.2 系列开源语言模型,3B、8B、30B 三种尺寸全部引入原生思维链推理(thinking),Apache 2.0 许可,可直接商用(官方发布:https:\u002F\u002Fresearch.ibm.com\u002Fblog\u002Fintroducing-granite-4-2)。这是一次针对企业 agent 场景的明确押注——模型不再只负责生成回复,而要在多步任务里规划、调用工具、自我纠错。\n\n## 原生思维链,做成了三档开关\n\nGranite 4.2 最值得玩味的产品化细节,是把\"思考\"变成了显式可切换的模式:thinking 模式下模型先输出完整的逐步推理再给答案;non-thinking 模式跳过推理直接回答;low-effort thinking 只做极简推理,内部独白可以短到一句\"Simple answer\"。三档切换通过 chat template 参数完成,不需要换模型。\n\n工具调用也与推理集成:模型会先想清楚\"该调哪个工具、为什么调\",再发起调用,工具定义沿用 OpenAI function schema。能在终端环境里导航代码库、处理多步开发任务的软件工程 agent,是这次训练的重点目标之一。\n\n## 训练流水线:多阶段 RL + 1T 合成代码\n\n按 IBM Research 的说法,这些能力不来自堆参数,而来自重新设计的训练流程。Granite 4.2 在 Granite 4.1 基座上做后训练:先过监督微调,再进入多阶段强化学习。第一段\"foundational RL\"覆盖全部三档尺寸,强化数学、科学、编码、推理和工具调用,混合可验证奖励与奖励模型评估;8B 和 30B 额外追加\"agentic RL\"阶段,专攻软件工程、终端编码和搜索驱动的工作流,最后叠加 RLHF 对齐。\n\n两个关键补充:一是用 IBM CodeAlchemy 管线生成的 1 万亿 token 合成代码参与训练;二是引入 mid-training 中间训练步骤以解锁更多推理能力。推理侧加了 speculative decoding 层,输出更快、并发更高,直接压低企业服务成本。IBM 还与 Hirundo 合作,用机器遗忘(machine unlearning)技术在不重训的前提下削减不良输出。\n\n## Benchmark:30B 拿下 SWE Bench Pro 同级最高分\n\n评测覆盖 AIME25(推理)、LiveCodeBench v6(编码)、IFBench 与 τ³-bench(指令遵循)、BFCL v4(工具调用)、SWE Bench Pro 和 Terminal-Bench 2.1(agent)。结果分层很清晰:3B 档在所有任务上明显领先同级;8B 档在推理和指令遵循上持平或超过对手,且是同尺寸里唯一报告 SWE Bench Pro 成绩的模型;30B 档拿下 SWE Bench Pro 最高分,面对同级或更大模型保持全面竞争力。\n\n## 顺带发布的语音模型,470M 跑到 RTFx 12,600\n\n同场发布的还有 Granite Speech 5.0 Turbo CTC 与非商业版 CTC NC:470M 参数、去掉 LLM backbone 的 CTC 架构,面向笔记本、手机等边缘设备。官方测试中其在单张 H200 上 RTFx 约 12,600,而 Hugging Face Open ASR 榜单的速度领先者约 6,000——吞吐翻倍,一秒可转录三小时录音,适合呼叫中心这类高音量转录场景。\n\n## 所以呢\n\n在全网都在卷万亿参数 MoE 的 2026 年,IBM 用 dense 架构 + 三档推理开关 + Apache 2.0 给出了一条差异化路线:不追绝对智能上限,而是把\"够聪明、可审计、可私有化部署\"做到企业敢用的程度。对要做私有 agent 的团队,一个 3B 全面领先同级、8B 跑完整 agent 评测、30B 敢和更大模型掰手腕的全开源家族,值得进选型清单。毕竟在企业市场,许可证和部署自由度,往往比榜单第一更有说服力。","https:\u002F\u002Fresearch.ibm.com\u002Fblog\u002Fintroducing-granite-4-2","653dda08-2edc-4d17-aeb2-56b0c88dd918",[11,15,18,21],{"id":12,"name":13,"slug":13,"description":14,"color":14},"6ad31a14-c0da-42df-81fd-564281f768db","agentic-ai",null,{"id":16,"name":17,"slug":17,"description":14,"color":14},"01598627-1ea6-4b27-a5d8-874971571a71","llm",{"id":19,"name":20,"slug":20,"description":14,"color":14},"7e89b5cc-57db-4f37-bc6d-28919a73931c","model-release",{"id":22,"name":23,"slug":23,"description":14,"color":14},"b9bd9039-fcdb-41a8-b85b-fc1587def2b9","open-source",[25],{"id":26,"lang":27,"title":28,"summary":29,"content":30},"2a115d3a-91fc-4c75-904f-aa207a3c8ffc","en","IBM Granite 4.2 Goes Open Source: Native Chain-of-Thought as a Switch, 30B Tops SWE Bench Pro","IBM releases Granite 4.2, an open-weights LLM family in 3B\u002F8B\u002F30B dense sizes with native chain-of-thought reasoning and a three-level thinking switch. The 30B model posts the top SWE Bench Pro score in its class, all under Apache 2.0.","As the frontier of large language models shifts from \"chatting\" to \"getting work done,\" IBM has delivered its answer: the Granite 4.2 family of open language models, released in 3B, 8B, and 30B sizes, all featuring native chain-of-thought reasoning (\"thinking\"), under an Apache 2.0 license that allows direct commercial use (official announcement: https:\u002F\u002Fresearch.ibm.com\u002Fblog\u002Fintroducing-granite-4-2). It is a clear bet on enterprise agent scenarios — the model is no longer just generating replies, but planning, calling tools, and self-correcting inside multi-step tasks.\n\n## Native chain-of-thought, built as a three-way switch\n\nThe most interesting product decision in Granite 4.2 is making \"thinking\" an explicitly switchable mode: in thinking mode the model outputs full step-by-step reasoning before the answer; non-thinking mode skips reasoning and answers directly; low-effort thinking performs minimal reasoning, with an internal monologue as short as a single \"Simple answer.\" All three levels are toggled via chat template parameters — no model swap required.\n\nTool calling is integrated with reasoning as well: the model first reasons about which tool to call and why, then makes the call, with tools defined using the OpenAI function schema. Software engineering agents that can navigate codebases, handle multi-step development tasks, and operate in terminal environments were a primary training target.\n\n## Training pipeline: multi-stage RL + 1T tokens of synthetic code\n\nAccording to IBM Research, these capabilities come not from parameter scale but from a redesigned training process. Granite 4.2 is post-trained on top of Granite 4.1 base models: supervised fine-tuning first, then multiple reinforcement learning phases. The first stage, \"foundational RL,\" was applied to all three sizes, strengthening math, science, coding, reasoning, and tool calling, combining verifiable rewards with reward-model-based evaluation. The 8B and 30B models continue with a specialized \"agentic RL\" phase focused on enterprise-style tasks — software engineering, terminal-based coding, and search-driven workflows — followed by RLHF alignment.\n\nTwo further ingredients: training incorporated 1 trillion tokens of synthetic code generated with IBM's CodeAlchemy pipeline, and an intermediate \"mid-training\" step shown to unlock additional reasoning power. On the inference side, a speculative decoding layer outputs text faster while serving more users, directly cutting enterprise serving costs. IBM is also working with Hirundo to apply machine unlearning to reduce undesirable outputs without fully retraining the models.\n\n## Benchmarks: the 30B takes the top SWE Bench Pro score in its class\n\nEvaluations cover AIME25 (reasoning), LiveCodeBench v6 (coding), IFBench and τ³-bench (instruction following), BFCL v4 (tool calling), plus SWE Bench Pro and Terminal-Bench 2.1 (agentic). The results are clearly stratified: at 3B, Granite leads decisively across all tasks; at 8B, it matches or exceeds competitors on reasoning and instruction following and is the only model in its size class reporting SWE Bench Pro results; at 30B, Granite achieves the highest SWE Bench Pro score and stays competitive across all benchmarks against similar or larger models.\n\n## A bonus speech model: 470M parameters hitting RTFx 12,600\n\nReleased alongside are Granite Speech 5.0 Turbo CTC and a non-commercial CTC NC variant: 470-million-parameter CTC models with no LLM backbone, aimed at laptops, smartphones, and other edge devices. In IBM's testing the Turbo CTC model reached roughly 12,600 RTFx on a single H200 GPU, versus around 6,000 for the speed leaders on the Hugging Face Open ASR leaderboard — twice the throughput, transcribing three hours of audio in about a second, well suited to high-volume workloads like call-center analytics.\n\n## So what\n\nIn a 2026 where everyone races toward trillion-parameter MoEs, IBM's differentiated route — dense architecture plus a three-level reasoning switch plus Apache 2.0 — doesn't chase the absolute intelligence ceiling; it makes models \"smart enough, auditable, and privately deployable\" to a degree enterprises will actually trust. For teams building private agents, a fully open family where 3B leads its tier across the board, 8B runs the full agent evaluation suite, and 30B goes toe-to-toe with larger models deserves a spot on the shortlist. In the enterprise market, licensing and deployment freedom often speak louder than a leaderboard crown.","ibm-granite-4-2-native-reasoning-agents","2026-08-28T14:00:00Z","2026-08-27T23:11:14.727452Z","2026-08-27T23:11:14.727462Z",true,"agent",209,[39,48],{"slug":40,"tag_slug":40,"title_zh":41,"title_en":42,"intro_zh":43,"intro_en":44,"id":45,"is_active":35,"created_at":46,"modified_at":47},"ai-for-science","AI for Science 2026：从 UniPert 到 GPT-Rosalind 的硬核进化","AI for Science 2026: from UniPert to GPT-Rosalind","生命科学、化学材料、物理世界模型——AI 正在从\"语言工具\"变成\"实验伙伴\"。本专题收录 AI 在三大科学方向的关键节点：UniPert 统一基因与化学扰动空间、GPT-Rosalind 端到端生命科学推理、达摩院 AI 智能体 28 小时找到 4 种超导新材料、Anthropic Claude Science 把工作台做成标准品。","From language tool to lab partner — AI is reshaping life sciences, chemistry\u002Fmaterials, and physical world models. This topic covers the key milestones: UniPert unifying genetic-chemical perturbation spaces, GPT-Rosalind's end-to-end life-sciences reasoning, DAMO's AI agent discovering 4 superconducting materials in 28 hours, and Anthropic's Claude Science workbench going mainstream.","988a4300-5fab-41c4-b5d8-63711a2dc757","2026-09-10T01:34:15.296649Z","2026-09-10T01:34:15.296663Z",{"slug":49,"tag_slug":49,"title_zh":50,"title_en":51,"intro_zh":52,"intro_en":53,"id":54,"is_active":35,"created_at":55,"modified_at":56},"h3-series","MiniMax H3 系列：从开源权重到 35 倍吞吐","MiniMax H3 Series: from open weights to 35x throughput","MiniMax H3 自 2026 年 8 月开源以来节奏密集：官方把生成、参考与编辑收回一个模型；ComfyUI 当天压进 RTX 3060；摩尔线程 3 小时完成国产 GPU 适配；fal 后训练版把吞吐拉到 35 倍；FastH3 蒸馏再砍推理成本。本专题持续追踪 H3 的发布—开源—蒸馏—部署全链路。","Since MiniMax open-sourced H3 in August 2026 the pace has been relentless: one unified omni-modal model, same-day ComfyUI support down to an RTX 3060, a 3-hour Day-0 port to Moore Threads GPUs, fal's post-trained H3 Max at 35x throughput, and FastH3 distillation cutting inference cost further. This topic tracks the full H3 chain — release, open weights, distillation, deployment.","83ef0daa-3c31-4cb3-86ed-e5ee58654d5f","2026-09-08T07:33:19.942193Z","2026-09-08T07:33:19.942209Z",{"items":58},[59,64,69,74,79,84],{"id":60,"title":61,"news_slug":62,"published_at":63},"f26ace13-9c96-47ea-a528-b6682a22aa1e","Apodex 1.1 把推理搬进真实执行:PIVOT-RL 定位关键决策点,35B mini 开源","apodex-1-1-agentic-execution-pivot-rl","2026-08-25T14:30:00+00:00",{"id":65,"title":66,"news_slug":67,"published_at":68},"b4754043-6b19-499f-8459-f8fc786f4d80","Pokee-Isaac 28B 把 10M 上下文塞进客户边界:28B 参数在 RULER 10M 上 93.3%","pokee-isaac-28b-10m-context","2026-08-20T14:00:00+00:00",{"id":70,"title":71,"news_slug":72,"published_at":73},"36055e5f-136f-497d-8763-3ed6609f59ff","Meta Muse Glimmer 30B 本地落地:Apache 2.0 的开源智能体,把 Agent 装进 24GB 显存","meta-muse-glimmer-30b-local-agent-apache2-r2","2026-08-19T03:00:00+00:00",{"id":75,"title":76,"news_slug":77,"published_at":78},"5878a668-282c-4b88-b2b8-7eef40b7938c","LFM2.5-2.6B：2.5GB 内存跑本机 Agent 220 tok\u002Fs","lfm2-5-2-6b-on-device-agent","2026-08-11T00:00:00+00:00",{"id":80,"title":81,"news_slug":82,"published_at":83},"79c1684f-f61d-4799-b3d0-6450c4ad10e8","Muse Glimmer:Meta 把 30B 「常驻本地的智能体」开源,把 Agent 拉到笔记本里 7×24 跑","muse-glimmer-30b-open-agentic-local","2026-08-10T00:00:00+00:00",{"id":85,"title":86,"news_slug":87,"published_at":88},"c94766df-827e-4e4e-a006-b6639ec76722","DeepSeek V4-Flash-0731 转正观察:权重不动,后训练把 Agent 分数打到 V4-Pro 之上","deepseek-v4-flash-0731-agent-benchmark-official-aug2026","2026-08-01T02:00:00+00:00"]