[{"data":1,"prerenderedAt":-1},["ShallowReactive",2],{"news-slug-gemma-4-26b-moe-3-8b-active-256k-apache2":3,"topics-all":36,"news-related-91f81ab2-f8c9-47a1-8919-3165d03f44b0":55},{"id":4,"title":5,"summary":6,"content":6,"original_url":7,"source_id":8,"tags":9,"translations":23,"news_slug":29,"published_at":30,"created_at":31,"modified_at":32,"is_published":33,"publish_type":34,"image_url":13,"view_count":35},"91f81ab2-f8c9-47a1-8919-3165d03f44b0","Gemma 4 26B：开源MoE模型的性价比新标杆","Google于2026年4月2日发布的Gemma 4模型家族中，26B MoE版本正在成为开源社区最受欢迎的选择。这个拥有260亿总参数、仅38亿激活参数的中型模型，在Apache 2.0许可下实现了前所未有的性价比突破。\n\nGemma 4 26B采用稀疏Mixture-of-Experts架构，每次前向传播只激活3.8B参数。这意味着在Q4量化后只需8GB显存即可运行——相当于一台普通笔记本的负载，却能达到接近GPT-4级别的推理能力。在MMLU基准测试中，它以83.2%的得分超越了Llama 4 Scout的79.8%和Qwen 3.5 Plus的82.1%。\n\n混合注意力机制是另一个亮点。Gemma 4 26B交替使用局部滑动窗口注意力和全局注意力，最后一层始终保持全局感知，使256K token的上下文窗口真正可用。这对于分析长代码仓库或整本技术文档尤为重要。\n\n全家族统一支持文本和图像多模态，E4B版本还额外支持音频输入。从树莓派到单块H100 GPU，Gemma 4覆盖了从边缘设备到数据中心的完整场景，这种「一个架构、多档硬件」的策略正在重新定义开源模型的部署边界。\n\n笔者认为，Gemma 4 26B的成功在于它找到了模型能力与推理成本的黄金分割点。当行业从「越大越好」转向「越精越好」，中型MoE模型很可能是下一代开源大模型的事实标准。","https:\u002F\u002Fwww.aimadetools.com\u002Fblog\u002Fgemma-4-family-guide\u002F","bd22b0c2-856a-4ce3-abf7-d2f644092c83",[10,14,17,20],{"id":11,"name":12,"slug":12,"description":13,"color":13},"0ef8513a-0a26-42f0-b6f9-5b6dadded45c","efficiency",null,{"id":15,"name":16,"slug":16,"description":13,"color":13},"8cf7490f-2449-4ba7-be19-61befa0d92b4","google",{"id":18,"name":19,"slug":19,"description":13,"color":13},"01598627-1ea6-4b27-a5d8-874971571a71","llm",{"id":21,"name":22,"slug":22,"description":13,"color":13},"b9bd9039-fcdb-41a8-b85b-fc1587def2b9","open-source",[24],{"id":25,"lang":26,"title":27,"summary":28,"content":13},"4071dfa1-ad71-4841-b1f8-3738d9af34db","en","Gemma 4 26B: Open-Source MoE Model's New Cost-Performance Benchmark","Among the Gemma 4 model family released by Google on April 2, 2026, the 26B MoE version is becoming the most popular choice in the open-source community. This mid-size model with 26B total parameters and only 3.8B activated parameters has achieved an unprecedented cost-performance breakthrough under Apache 2.0 license.\n\nGemma 4 26B adopts a sparse Mixture-of-Experts architecture, activating only 3.8B parameters per forward pass. This means after Q4 quantization, it only needs 8GB VRAM to run — equivalent to a regular laptop's load, yet achieves reasoning capability close to GPT-4 level. On the MMLU benchmark, it scored 83.2%, surpassing Llama 4 Scout's 79.8% and Qwen 3.5 Plus's 82.1%.\n\nThe hybrid attention mechanism is another highlight. Gemma 4 26B alternates between local sliding-window attention and global attention, with the last layer always maintaining global perception, making the 256K token context window truly usable. This is especially important for analyzing long code repositories or entire technical documentation.\n\nThe whole family uniformly supports text and image multimodality, with the E4B version additionally supporting audio input. From Raspberry Pi to single H100 GPU, Gemma 4 covers the full spectrum from edge devices to data centers, with this \"one architecture, multiple hardware tiers\" strategy redefining the deployment boundary of open-source models.\n\nThe author believes Gemma 4 26B's success lies in finding the golden ratio between model capability and inference cost. As the industry shifts from \"the bigger the better\" to \"the more refined the better,\" mid-size MoE models are likely to be the de facto standard for the next generation of open-source LLMs.","gemma-4-26b-moe-3-8b-active-256k-apache2","2026-04-26T19:00:00Z","2026-04-26T19:08:14.180842Z","2026-08-19T02:08:40.142862Z",true,"agent",164,[37,46],{"slug":38,"tag_slug":38,"title_zh":39,"title_en":40,"intro_zh":41,"intro_en":42,"id":43,"is_active":33,"created_at":44,"modified_at":45},"ai-for-science","AI for Science 2026：从 UniPert 到 GPT-Rosalind 的硬核进化","AI for Science 2026: from UniPert to GPT-Rosalind","生命科学、化学材料、物理世界模型——AI 正在从\"语言工具\"变成\"实验伙伴\"。本专题收录 AI 在三大科学方向的关键节点：UniPert 统一基因与化学扰动空间、GPT-Rosalind 端到端生命科学推理、达摩院 AI 智能体 28 小时找到 4 种超导新材料、Anthropic Claude Science 把工作台做成标准品。","From language tool to lab partner — AI is reshaping life sciences, chemistry\u002Fmaterials, and physical world models. This topic covers the key milestones: UniPert unifying genetic-chemical perturbation spaces, GPT-Rosalind's end-to-end life-sciences reasoning, DAMO's AI agent discovering 4 superconducting materials in 28 hours, and Anthropic's Claude Science workbench going mainstream.","988a4300-5fab-41c4-b5d8-63711a2dc757","2026-09-10T01:34:15.296649Z","2026-09-10T01:34:15.296663Z",{"slug":47,"tag_slug":47,"title_zh":48,"title_en":49,"intro_zh":50,"intro_en":51,"id":52,"is_active":33,"created_at":53,"modified_at":54},"h3-series","MiniMax H3 系列：从开源权重到 35 倍吞吐","MiniMax H3 Series: from open weights to 35x throughput","MiniMax H3 自 2026 年 8 月开源以来节奏密集：官方把生成、参考与编辑收回一个模型；ComfyUI 当天压进 RTX 3060；摩尔线程 3 小时完成国产 GPU 适配；fal 后训练版把吞吐拉到 35 倍；FastH3 蒸馏再砍推理成本。本专题持续追踪 H3 的发布—开源—蒸馏—部署全链路。","Since MiniMax open-sourced H3 in August 2026 the pace has been relentless: one unified omni-modal model, same-day ComfyUI support down to an RTX 3060, a 3-hour Day-0 port to Moore Threads GPUs, fal's post-trained H3 Max at 35x throughput, and FastH3 distillation cutting inference cost further. This topic tracks the full H3 chain — release, open weights, distillation, deployment.","83ef0daa-3c31-4cb3-86ed-e5ee58654d5f","2026-09-08T07:33:19.942193Z","2026-09-08T07:33:19.942209Z",{"items":56},[57,62,67,72,77,82],{"id":58,"title":59,"news_slug":60,"published_at":61},"98bece1e-3ea5-4ee6-b4b7-b4e9aa385fbc","Gemma 4为何变快了：Google多Token预测让本地推理提速3倍","gemma-4-mtp-drafter-74m-3x-local","2026-05-06T19:00:00+00:00",{"id":63,"title":64,"news_slug":65,"published_at":66},"68812025-96eb-4ca9-a1bc-8a82a40174dc","Google RRSI:给 Agent 外壳自进化加正则化","google-rrsi-agent-harness-regularization","2026-09-22T23:08:34+00:00",{"id":68,"title":69,"news_slug":70,"published_at":71},"3559e613-9558-48e1-ab20-f53b62796363","让每个 token 用上全部专家:高德 IntBMoE 解耦参与度、计算与显存,60ms 服务数亿用户","intbmoe-full-participation-block-moe","2026-09-21T13:01:54+00:00",{"id":73,"title":74,"news_slug":75,"published_at":76},"28c41f06-d20f-481c-b133-cd109af3aed1","答对之后停不下来:微软团队揪出在线蒸馏的 EOS 错配元凶","eos-mismatch-opd-length-inflation","2026-09-18T21:09:06+00:00",{"id":78,"title":79,"news_slug":80,"published_at":81},"2e27016d-b90e-45c7-825a-41fd1e435c80","JHU 新研究:组合持续学习机制,百任务记忆留存从 1.2% 提到 34.9%","compose-cl-long-horizon-memorization","2026-09-16T15:10:00+00:00",{"id":83,"title":84,"news_slug":85,"published_at":86},"30fca629-bace-4832-9789-b44aa8c8989d","学生团队从零训出开源 7B 模型 ZGCM-1:数学推理硬刚 235B 前沿","zgcm-1-open-7b-foundation-model","2026-09-15T19:10:00+00:00"]