[{"data":1,"prerenderedAt":-1},["ShallowReactive",2],{"news-slug-inkling-small-multimodal-moe-efficiency":3,"news-related-804b44fb-66c6-4355-83a4-b3a03a776d2a":38},{"id":4,"title":5,"summary":6,"content":7,"original_url":8,"source_id":9,"tags":10,"translations":24,"news_slug":31,"published_at":32,"created_at":33,"modified_at":34,"is_published":35,"publish_type":36,"image_url":14,"view_count":37},"804b44fb-66c6-4355-83a4-b3a03a776d2a","Inkling-Small 开放权重：12B 激活参数换来更高 Agent 效率，也暴露事实性短板","Thinking Machines Lab 发布开放权重多模态 MoE 模型 Inkling-Small：总参数 276B、每 token 激活 12B，原生处理文本、图像和音频，并支持最高 100 万 token 上下文。官方评测显示它在编码与 Agent 任务上超过更大的 Inkling，但事实问答明显退步；NVFP4 检查点将部署门槛降至 180GB 聚合显存。","# Inkling-Small：不是把模型砍小，而是重新训练出更高效的多模态 MoE\n\nThinking Machines Lab 在 7 月 30 日发布了开放权重模型 **Inkling-Small**。名字里虽然有 “Small”，但它仍有 276B 总参数，只是每个 token 激活 12B；相比此前 Inkling 的 975B 总参数、41B 激活参数，规模约为四分之一。模型采用 Apache 2.0 许可，权重已放到 Hugging Face，并可通过 Tinker 体验和微调。[官方发布页](https:\u002F\u002Fthinkingmachines.ai\u002Fnews\u002Finkling-small\u002F) 给出的重点不是“更小”，而是用更少计算维持接近甚至更好的推理与 Agent 能力。\n\n## 它为什么能缩到四分之一\n\nInkling-Small 是一个 42 层、仅解码器的稀疏 MoE Transformer。每个 token 会路由到 256 个专家中的 6 个，同时还有 2 个共享专家持续参与。它原生接收文本、图像和音频，三种输入都被投射到共享隐藏空间，再由同一个解码器联合处理；上下文窗口最高 100 万 token，输出仍为文本。\n\n这次缩小并非简单蒸馏旧模型。官方说明，Inkling-Small 启动训练晚于大模型，因此团队调整了预训练数据组合和训练方法。其早期预览检查点一部分使用 Inkling 作为教师做在线策略蒸馏，随后又继续进行了两周的 Agent 编码强化学习。换句话说，它是在更小的激活预算下重新优化训练路线，而不是把大模型机械压缩一遍。\n\n## 编码与推理更强，但事实性明显退步\n\n官方评测显示，Inkling-Small 在 SWE-bench Verified 上取得 80.2%，高于 Inkling 的 77.6%；Terminal-Bench 2.1 为 64.7%，略高于 63.8%；纯文本 Humanity’s Last Exam 为 31.6%，也高于 29.7%。这些数字支持团队的判断：较小模型在推理、编码和工具使用上可以超过教师模型。\n\n但代价同样清楚。SimpleQA Verified 从 Inkling 的 43.9% 降到 20.6%，Tau 3 Banking 从 23.7% 降到 15.5%。官方因此明确写道，Inkling 仍然在知识覆盖和事实性上占优。这组反差比单看榜单更重要：强化学习和蒸馏可以把有限计算集中到“会做任务”，却不自动补回知识覆盖。\n\n## 真正的部署门槛仍然很高\n\nInkling-Small 提供 BF16 和 NVFP4 等数值格式。官方模型卡显示，BF16 检查点至少需要 600GB 聚合显存，可由 4 张 NVIDIA B300 或 8 张 H200 承载；NVFP4 将要求降到至少 180GB，可在 1 张 B300 上以 W4A4 运行，或在 2 张 H200 上以 W4A16 运行。支持的部署框架包括 SGLang、vLLM、TokenSpeed、Unsloth 和 Hugging Face。\n\n所以，这不是普通消费级显卡能本地跑的“小模型”。它的意义，是把一个原生文本、图像、音频模型从超大集群门槛，压到单张数据中心级 GPU 或双卡 H200 的范围。对需要私有部署、工具调用和多模态输入的团队，这种变化比“总参数少了多少”更实际。\n\nInkling-Small 给出的信号很直接：开放权重模型的下一轮竞争，不只是继续堆总参数，而是把激活参数、推理预算、量化格式和训练配方一起优化。真正有价值的“小”，不是名字更小，而是单位计算能完成更多任务——同时把它不擅长的事实检索，明确交给检索与校验系统。","https:\u002F\u002Fthinkingmachines.ai\u002Fnews\u002Finkling-small\u002F","95239a8d-29f2-486d-84ca-28174cab2405",[11,15,18,21],{"id":12,"name":13,"slug":13,"description":14,"color":14},"01598627-1ea6-4b27-a5d8-874971571a71","llm",null,{"id":16,"name":17,"slug":17,"description":14,"color":14},"7e89b5cc-57db-4f37-bc6d-28919a73931c","model-release",{"id":19,"name":20,"slug":20,"description":14,"color":14},"499f4b56-819d-49a3-9609-33e775143b86","multimodal",{"id":22,"name":23,"slug":23,"description":14,"color":14},"b9bd9039-fcdb-41a8-b85b-fc1587def2b9","open-source",[25],{"id":26,"lang":27,"title":28,"summary":29,"content":30},"b58b8895-d054-4b71-a577-2b19706042e3","en","Inkling-Small: 12B-active weights trade facts for efficiency","Thinking Machines Lab has open-released Inkling-Small, a 276B-total \u002F 12B-active multimodal MoE under Apache 2.0. Native text, image and audio inputs, up to 1M-token context. Official evals show it edges past the larger Inkling on coding and reasoning, but loses more than half its SimpleQA Verified score. NVFP4 checkpoint cuts the deployment floor to 180 GB of aggregated VRAM.","# Inkling-Small Is Not Just a Compressed Model: It Reworks the Multimodal MoE Efficiency Trade-Off\n\nThinking Machines Lab released the open-weight **Inkling-Small** model on July 30. Despite the word “Small,” it still contains 276 billion total parameters, with 12 billion activated for each token. The earlier Inkling model has 975 billion total parameters and 41 billion active parameters, making the new model roughly one quarter of its size by both measures. Inkling-Small is licensed under Apache 2.0, its weights are available on Hugging Face, and it can also be tested and fine-tuned through Tinker. The central claim in the [official release](https:\u002F\u002Fthinkingmachines.ai\u002Fnews\u002Finkling-small\u002F) is not merely that the model is smaller, but that it preserves or improves reasoning and agentic performance while using substantially less compute.\n\n## How the model reduces its active footprint\n\nInkling-Small is a 42-layer, decoder-only Transformer with a sparse Mixture-of-Experts feed-forward backbone. For every token, the router selects 6 of 256 experts, while 2 shared experts remain active for all tokens. The model accepts text, image, and audio inputs natively. Each modality is projected into a shared hidden space and processed jointly by the same decoder. It supports a context window of up to one million tokens and produces text output.\n\nThe smaller design is not simply a mechanical compression of the older model. Thinking Machines Lab says Inkling-Small began training after Inkling, which allowed the team to revise both the pre-training data mixture and the machine-learning recipe. An earlier preview checkpoint was post-trained partly through on-policy distillation, with Inkling acting as the teacher. The team then continued scaling agentic coding reinforcement learning for two additional weeks. In other words, the company redesigned the training path around a lower active-parameter budget rather than merely shrinking the original checkpoint.\n\n## Better coding and reasoning, but a sharp factuality regression\n\nThe official evaluation suite reports 80.2% on SWE-bench Verified for Inkling-Small, compared with 77.6% for Inkling. On Terminal-Bench 2.1, the scores are 64.7% and 63.8%, respectively. On the text-only version of Humanity’s Last Exam, Inkling-Small reaches 31.6%, ahead of Inkling at 29.7%. These results support the lab’s conclusion that the smaller model can outperform its teacher on reasoning, coding, and tool-use tasks.\n\nThe trade-off is equally visible. On SimpleQA Verified, Inkling-Small drops to 20.6% from Inkling’s 43.9%. On Tau 3 Banking, it scores 15.5%, compared with 23.7% for the larger model. Thinking Machines Lab explicitly states that Inkling retains an advantage in knowledge coverage and factuality. That contrast matters more than a single headline benchmark: reinforcement learning and distillation can concentrate limited compute on completing tasks, but they do not automatically restore broad factual knowledge.\n\n## The deployment threshold is lower, not low\n\nInkling-Small is distributed in BF16 and NVFP4 formats, among others listed in the model card. The BF16 checkpoint requires at least 600 GB of aggregated VRAM, which can be provided by four NVIDIA B300 GPUs or eight H200 GPUs. The NVFP4 checkpoint lowers the requirement to at least 180 GB. It can run in W4A4 mode on one B300, subject to the SM100-or-newer requirement, or in W4A16 mode on two H200 GPUs. Supported deployment frameworks include SGLang, vLLM, TokenSpeed, Unsloth, and Hugging Face.\n\nThis is therefore not a “small model” that runs on an ordinary consumer GPU. Its practical significance is that a native text-image-audio model moves from a very large cluster requirement to a single data-center GPU or a pair of H200s. For teams that need private deployment, tool use, and multimodal input, that reduction is more meaningful than the label attached to the parameter count.\n\nInkling-Small points to a broader direction for open-weight models. The next stage of competition is not only about increasing total parameters. It is about jointly optimizing activated parameters, reasoning effort, quantization formats, and training recipes. A useful small model is not one with a smaller name; it is one that completes more work per unit of compute while clearly handing its factual weaknesses to retrieval and verification systems.","inkling-small-multimodal-moe-efficiency","2026-08-05T16:32:13Z","2026-08-05T16:47:28.592135Z","2026-08-05T16:47:28.592143Z",true,"agent",102,{"items":39},[40,45,50,55,60,65],{"id":41,"title":42,"news_slug":43,"published_at":44},"40095b51-97b0-4fd4-9b1d-f636c970572e","阿里 Qwen 团队发布 Qwen3.8-Max:2.4 万亿参数 MoE 模型首度开放权重","qwen3-8-max-2-4t-moe-open-weights","2026-08-07T02:00:00+00:00",{"id":46,"title":47,"news_slug":48,"published_at":49},"a151db0c-d832-4df2-ac03-2d4e58b26e99","Kimi K3 跑通 MiniTriton:Moonshot 让 LLM 第一次从零编译出自己的 GPU 编译器","kimi-k3-minitriton-gpu-compiler","2026-07-26T14:00:00+00:00",{"id":51,"title":52,"news_slug":53,"published_at":54},"dfc3dec4-2211-4c7e-b6ff-9e0d9a479ec4","微软与 Mistral 签下数十亿美元协议:Vera Rubin GPU 上的「欧洲主权云」开始落地","microsoft-mistral-vera-rubin-sovereign","2026-07-22T02:00:00+00:00",{"id":56,"title":57,"news_slug":58,"published_at":59},"cf01282f-8a64-49a8-a608-9b806ccfbea3","Mira Murati 实验室 Inkling 开源：975B MoE 不卷\"最强\"，押注\"可定制\"","thinking-machines-inkling","2026-07-15T22:00:00+00:00",{"id":61,"title":62,"news_slug":63,"published_at":64},"1a27bedc-012d-4d63-85e1-ddc57aabd8bf","ByteDance UniVR 让模型「在视觉空间里思考」：34B 参数逼近 Gemini 3 Pro + Nano Banana 2","bytedance-univr-34b","2026-07-14T12:10:00+00:00",{"id":66,"title":67,"news_slug":68,"published_at":69},"66079e92-3544-45b0-abeb-31d628220449","百度 Unlimited OCR：把端到端文档解析推进「一次性长文档」时代，R-SWA 把 KV 缓存压成常数","baidu-unlimited-ocr-rswa-constant-kv","2026-06-29T08:00:00+00:00"]