[{"data":1,"prerenderedAt":-1},["ShallowReactive",2],{"news-slug-nvidia-iterative-puzzle-nemotron-3-super":3,"news-related-d04db7ae-1027-4ba5-8939-563fd7372c21":36},{"id":4,"title":5,"summary":6,"content":6,"original_url":7,"source_id":8,"tags":9,"translations":23,"news_slug":29,"published_at":30,"created_at":31,"modified_at":32,"is_published":33,"publish_type":34,"image_url":13,"view_count":35},"d04db7ae-1027-4ba5-8939-563fd7372c21","Iterative Puzzle：Nemotron-3 砍到 62%，吞吐 2.03×","大模型部署成本的核心瓶颈从来不是参数量,而是单节点能撑多少并发请求——MoE 模型尤其如此,active parameters、KV cache 与 Mamba state 共同卡死了上限。\n\nNVIDIA Nemotron 团队发布 Nemotron-Labs-3-Puzzle-75B-A9B(arXiv: 2607.04371),把 120.7B\u002F12.8B active 的 Nemotron-3-Super 压到 75.3B\u002F9.3B active,却不是均匀剪枝。核心方法叫 Iterative Puzzle:把 MoE 中间通道、激活专家数、Mamba SSM state 一并扔进混合整数规划求解器,按部署 SLA 反向选每层最优实现。三阶段压缩各配 24B\u002F43.2B\u002F52.8B token 的 KD 恢复,长上下文阶段再扩到 128K–512K 微调。\n\n量化到 NVFP4 后,8xB200 服务吞吐 +2.03x(8K\u002F64K 解码),单卡 H100 1M 上下文并发从 1 涨到 8——权重从 70GB 压到 44.5GB。代价是 Arena-Hard-V2 -4.2、SWE-Bench -2.6(指令遵循与 agent 类损失最大);长上下文 RULER 1M 仅掉 1.7,几乎无损。\n\n这套以部署换性能的工作流,大概率会成为下一代开源 LLM 的标配——把 expert 中间维度、top-k、Mamba state、注意力层都放进同一个 NAS 求解器,正是当前社区仍欠缺的工程纪律。配合 NVFP4 与多 token 预测头,4090 和 H100 都能跑出旗舰吞吐,直接拉低开源大模型落地的算力门槛。","https:\u002F\u002Farxiv.org\u002Fabs\u002F2607.04371","7437aeb9-930c-4866-a2e9-48003c1a792b",[10,14,17,20],{"id":11,"name":12,"slug":12,"description":13,"color":13},"2d9c2fb0-2be5-4ad1-aedb-e9747addf355","compression",null,{"id":15,"name":16,"slug":16,"description":13,"color":13},"0ef8513a-0a26-42f0-b6f9-5b6dadded45c","efficiency",{"id":18,"name":19,"slug":19,"description":13,"color":13},"7e89b5cc-57db-4f37-bc6d-28919a73931c","model-release",{"id":21,"name":22,"slug":22,"description":13,"color":13},"8dac812d-3839-4abe-a855-5f56ec9515fd","nvidia",[24],{"id":25,"lang":26,"title":27,"summary":28,"content":28},"4353ecc8-e52f-40ce-ad9d-c66ebc3416dc","en","Iterative Puzzle trims Nemotron-3 to 62%, 2.03x throughput","The core bottleneck of large-model deployment cost has never been parameter count, but how many concurrent requests a single node can hold — and this is especially true for MoE models, where active parameters, KV cache and Mamba state together cap the upper limit. NVIDIA's Nemotron team released Nemotron-Labs-3-Puzzle-75B-A9B (arXiv: 2607.04371), compressing the 120.7B \u002F 12.8B-active Nemotron-3-Super down to 75.3B \u002F 9.3B-active, but with non-uniform pruning. The core method is called Iterative Puzzle: it feeds the MoE intermediate channels, the number of active experts, and the Mamba SSM state into a mixed-integer-programming solver, choosing the optimal per-layer implementation backwards from the deployment SLA. Three compression stages each have 24B\u002F43.2B\u002F52.8B tokens of KD recovery, with the long-context stage then expanded to 128K–512K fine-tuning. After NVFP4 quantization, 8×B200 service throughput is +2.03× (8K\u002F64K decoding), and single-card H100 1M-context concurrency goes from 1 to 8 — weights shrink from 70GB to 44.5GB. The cost is Arena-Hard-V2 -4.2, SWE-Bench -2.6 (instruction-following and agent-class losses the largest); long-context RULER 1M drops only 1.7, almost lossless. This deployment-for-performance workflow will most likely become the standard for the next generation of open-source LLMs — putting expert intermediate dimensions, top-k, Mamba state, and attention layers all into the same NAS solver is exactly the engineering discipline the community is still missing. Combined with NVFP4 and multi-token prediction heads, 4090 and H100 can both deliver flagship throughput, directly pulling down the compute threshold for open-source large-model landing.","nvidia-iterative-puzzle-nemotron-3-super","2026-07-18T06:00:00Z","2026-07-18T06:04:47.277509Z","2026-08-19T02:08:40.142862Z",true,"agent",81,{"items":37},[38,43,48,53,58,63],{"id":39,"title":40,"news_slug":41,"published_at":42},"41c8976c-d775-4603-aa01-693c57b7b0bd","NVIDIA Nemotron 3 Embed 登顶 RTEB：把 8B 旗舰检索能力蒸馏进 1B 部署款","nvidia-nemotron-3-embed","2026-07-16T18:00:00+00:00",{"id":44,"title":45,"news_slug":46,"published_at":47},"4f19e049-42e2-43f3-9a9e-4ff97cda00dd","CHERRY 用「15% 监督 + 6 层折叠 + 专家融合」三件套把 LLM 训练推到新性价比边界","cherry-selective-token-training","2026-07-01T22:00:00+00:00",{"id":49,"title":50,"news_slug":51,"published_at":52},"5e8f7048-3b9e-4a41-9aa4-de33e21f3609","Nemotron-Labs-TwoTower：AR\u002F扩散双塔解耦，吞吐 2.42×","nvidia-twotower-ar-diffusion","2026-07-01T20:00:00+00:00",{"id":54,"title":55,"news_slug":56,"published_at":57},"fb97a60d-69a1-4988-8de6-d1540ba63359","2.4B 参数读懂整页 A4:Cohere Labs 把最小的多模态模型挂上了 Apache 2.0","cohere-north-micro-vision-open-vlm","2026-08-18T13:30:00+00:00",{"id":59,"title":60,"news_slug":61,"published_at":62},"de2cceb2-7d39-4a5f-844e-5a3144667f49","Nemotron 3.5 Lightning 开源：30B 总参 3B 激活的混合 MoE，直接用 NVFP4 配方预训练","nemotron-35-lightning-30b-a3b-open-release","2026-08-16T15:00:00+00:00",{"id":64,"title":65,"news_slug":66,"published_at":67},"4244f57a-3afa-465c-aa67-793df6eba5cc","LFM2.5-VL-3B 开源：3.1B 参数让手机读懂屏幕、框住物体、自己调工具","liquid-ai-lfm2-5-vl-3b-edge-vlm","2026-08-14T13:30:00+00:00"]