[{"data":1,"prerenderedAt":-1},["ShallowReactive",2],{"news-slug-ibm-granite-4-1-dense-8b-moe-32b-grc":3,"news-related-3c9e4d7f-6f2a-4fea-8d22-c351b8fd7a4a":36},{"id":4,"title":5,"summary":6,"content":6,"original_url":7,"source_id":8,"tags":9,"translations":23,"news_slug":29,"published_at":30,"created_at":31,"modified_at":32,"is_published":33,"publish_type":34,"image_url":13,"view_count":35},"3c9e4d7f-6f2a-4fea-8d22-c351b8fd7a4a","IBM Granite 4.1：Dense架构回归，8B参数挑战32B MoE性能","2026年4月，LLM江湖的主角无疑是大规模MoE架构——Llama 4 Scout、DeepSeek V4、Qwen3.6系列，个个都是千亿参数起步。但IBM偏偏在这个节点发布了纯Dense路线的Granite 4.1。\n\nGranite 4.1是一个Dense decoder-only模型家族，提供3B、8B、30B三种规格。参数不大，训练规模并不敷衍：15T tokens、五阶段预训练流水线，其中第五阶段将上下文窗口阶段性扩展至512K，并采用含DAPO loss的四阶段RLHF。\n\n更值得关注的是8B版本的效率——它能匹配上一代32B MoE模型的性能，说明Dense架构并非天然低效，只要训练足够精良。30B版本可部署在单张H100上，对于需要私有化部署的企业用户，这个组合很有吸引力。\n\n真正的差异在于数据治理。IBM在预训练数据阶段就嵌入了GRC评估，这一步用户看不到，但对金融、医疗等受监管行业意义重大。\n\n不过，工程严谨性只是门槛，生产稳定性才是最终验证。","https:\u002F\u002Fhuggingface.co\u002Fblog\u002Fibm-granite\u002Fgranite-4-1","653dda08-2edc-4d17-aeb2-56b0c88dd918",[10,14,17,20],{"id":11,"name":12,"slug":12,"description":13,"color":13},"40269b40-7942-4650-9672-ed2e6524d37a","ai-technology",null,{"id":15,"name":16,"slug":16,"description":13,"color":13},"0ef8513a-0a26-42f0-b6f9-5b6dadded45c","efficiency",{"id":18,"name":19,"slug":19,"description":13,"color":13},"01598627-1ea6-4b27-a5d8-874971571a71","llm",{"id":21,"name":22,"slug":22,"description":13,"color":13},"b9bd9039-fcdb-41a8-b85b-fc1587def2b9","open-source",[24],{"id":25,"lang":26,"title":27,"summary":28,"content":13},"2a14897d-e49d-4189-ad3e-bbff928e4541","en","IBM Granite 4.1: dense is back, 8B challenges 32B MoE","In April 2026, the LLM arena's main characters are undoubtedly large-scale MoE architectures — Llama 4 Scout, DeepSeek V4, Qwen3.6 series, all starting at hundreds of billions of parameters. But at exactly this point, IBM released Granite 4.1, taking the pure Dense route.\n\nGranite 4.1 is a Dense decoder-only model family, available in 3B, 8B, and 30B sizes. The parameters aren't big, but the training scale isn't half-hearted either: 15T tokens, five-stage pretraining pipeline, where the fifth stage extends the context window in stages to 512K and adopts four-stage RLHF with DAPO loss.\n\nMore noteworthy is the 8B version's efficiency — it can match the previous-generation 32B MoE model's performance, showing that Dense architecture isn't inherently inefficient as long as training is sufficiently refined. The 30B version is deployable on a single H100, an attractive combination for enterprise users needing private deployment.\n\nThe real differentiator is data governance. IBM embeds GRC assessment at the pretraining data stage — users don't see this step, but it's significant for regulated industries like finance and healthcare.\n\nThat said, engineering rigor is only the threshold; production stability is the final validation.","ibm-granite-4-1-dense-8b-moe-32b-grc","2026-04-29T19:10:00Z","2026-04-29T19:09:20.688876Z","2026-08-19T02:08:40.142862Z",true,"agent",155,{"items":37},[38,43,48,53,58,63],{"id":39,"title":40,"news_slug":41,"published_at":42},"4a89fe5a-8703-49e5-b083-079cbda0fa2a","蒸馏也有副作用:中间训练期上KD,推理上涨、事实记忆反而变慢","switch-distillation-midtraining-kd","2026-09-02T17:10:00+00:00",{"id":44,"title":45,"news_slug":46,"published_at":47},"199cd4ef-f092-45a5-8635-91778dd2bce2","编译即训练：一句规约炼出 83.6% 准确率的本地神经函数，教师模型只用一次","compile-by-training-neural-functions","2026-09-04T23:08:03+00:00",{"id":49,"title":50,"news_slug":51,"published_at":52},"7a6d28b6-65da-4a29-96b1-dedb9894de97","随机驱逐追平最强打分器:Salesforce 重写 KV Cache 压缩常识","random-attention-kv-cache-eviction","2026-09-04T19:08:26+00:00",{"id":54,"title":55,"news_slug":56,"published_at":57},"6d60f9ba-4866-4520-9f12-d955e37f8472","Gated DeltaNet 全压 4-bit 没掉点:一篇论文拆掉混合 LLM 的量化禁忌","gated-deltanet-nvfp4-full-4bit","2026-09-04T15:08:02+00:00",{"id":59,"title":60,"news_slug":61,"published_at":62},"b1645fba-d364-47e6-97da-06868f98d987","Linux 内核 7.x 每版近 2000 个 CVE:AI 帮倒忙,维护者不堪重负","linux-kernel-cve-ai-overwhelmed","2026-09-04T00:00:00+00:00",{"id":64,"title":65,"news_slug":66,"published_at":67},"44a035c8-b8a3-48e5-af4f-c76323dac7b5","RWKV7-G1j 13.3B 开源:不用注意力,每 token 推理成本是常数","rwkv7-g1j-13b-attention-free","2026-09-03T13:14:19+00:00"]