[{"data":1,"prerenderedAt":-1},["ShallowReactive",2],{"news-slug-llama-4-scout-maverick-17b-active-10m-context":3,"news-related-3003735b-bbf9-44f4-80a7-d563efdce828":36},{"id":4,"title":5,"summary":6,"content":6,"original_url":7,"source_id":8,"tags":9,"translations":23,"news_slug":29,"published_at":30,"created_at":31,"modified_at":32,"is_published":33,"publish_type":34,"image_url":13,"view_count":35},"3003735b-bbf9-44f4-80a7-d563efdce828","Llama 4：Meta用MoE架构重新定义开源大模型效率边界","2026年4月，Meta发布Llama 4开源大模型家族，全面转向Mixture-of-Experts（MoE）混合专家架构，一改往日dense transformer路线。Llama 4系列在保持开源权重可获取的同时，在多项基准上逼近甚至超越GPT-4o、Gemini 2.0 Flash等闭源头部模型，被视为开源大模型史上最重要的一次架构升级。\n\nLlama 4家族包含两款主力模型：Llama 4 Scout与Llama 4 Maverick。Scout总参数量109B，每次推理仅激活17B参数（16位专家），支持高达1000万token超长上下文，可一次性处理整个代码库或整本书籍级别的任务。旗舰模型Maverick总参数400B，同样每次只激活17B（128位专家），在LMArena基准上突破1400分，超越GPT-4o和Gemini 2.0 Flash。\n\nMoE架构的核心逻辑是稀疏激活：并非每个token都经过全部400B参数计算，而是动态路由到最相关的专家子网络。一台8×H100 GPU节点即可跑出GPT-4级别质量，推理成本降至闭源模型的五分之一左右。Scout在Int4量化后甚至可单卡H100运行，大幅降低本地部署门槛。\n\n开源权重意味着可自由下载、量化和fine-tune。4月以来，Together AI、Fireworks AI等主流推理平台均已上线Llama 4 API，Ollama也支持本地一键拉取。对受限于预算或数据隐私的团队，Llama 4 Maverick提供了可比较的能力同时成本大幅降低，这本身就是一次效率革命。\n\n从技术演进看，Llama 4验证了MoE在超大规模开源模型上的可行性。可以预见，稀疏激活将成为开源大模型的主流方向。","https:\u002F\u002Ffazm.ai\u002Fblog\u002Fllm-model-release-april-2026","c9bf74a9-b202-46b4-b3de-33466d933bfa",[10,14,17,20],{"id":11,"name":12,"slug":12,"description":13,"color":13},"5e628969-6d2a-437f-998a-104e4b16cfb1","ai-progress",null,{"id":15,"name":16,"slug":16,"description":13,"color":13},"01598627-1ea6-4b27-a5d8-874971571a71","llm",{"id":18,"name":19,"slug":19,"description":13,"color":13},"b1853a5a-d940-42b7-94f9-0488ee3f2cf7","new-model",{"id":21,"name":22,"slug":22,"description":13,"color":13},"b9bd9039-fcdb-41a8-b85b-fc1587def2b9","open-source",[24],{"id":25,"lang":26,"title":27,"summary":28,"content":13},"3ff80e30-c867-47e2-b5b9-6505e07574fa","en","Llama 4: Meta redefines open-source efficiency with MoE","In April 2026, Meta released the Llama 4 open-source LLM family, fully transitioning to the Mixture-of-Experts (MoE) hybrid-expert architecture, a change from the previous Dense Transformer route. The Llama 4 series, while maintaining open-source weight availability, approaches or even surpasses closed-source top-tier models like GPT-4o and Gemini 2.0 Flash on multiple benchmarks, seen as the most important architectural upgrade in open-source LLM history.\n\nThe Llama 4 family includes two flagship models: Llama 4 Scout and Llama 4 Maverick. Scout has 109B total parameters, activating only 17B per inference (16 experts), supporting up to 10 million tokens of ultra-long context, able to process entire codebases or book-level tasks at once. The flagship Maverick has 400B total parameters, also activating only 17B per inference (128 experts), breaking 1400 on the LMArena benchmark, surpassing GPT-4o and Gemini 2.0 Flash.\n\nThe core logic of MoE architecture is sparse activation: not every token goes through all 400B parameters of computation, but is dynamically routed to the most relevant expert subnetwork. A single 8×H100 GPU node can deliver GPT-4-level quality, with inference cost dropping to about one-fifth of closed-source models. Scout after Int4 quantization can even run on a single H100, greatly lowering the local deployment threshold.\n\nOpen-source weights mean free download, quantization, and fine-tuning. Since April, major inference platforms including Together AI and Fireworks AI have launched Llama 4 APIs, and Ollama also supports local one-click pull. For teams constrained by budget or data privacy, Llama 4 Maverick offers comparable capability with significantly lower cost — an efficiency revolution in itself.\n\nFrom a technical evolution perspective, Llama 4 validates the feasibility of MoE on ultra-large-scale open-source models. It's foreseeable that sparse activation will become the mainstream direction for open-source LLMs.","llama-4-scout-maverick-17b-active-10m-context","2026-04-26T10:10:00Z","2026-04-26T10:07:42.408910Z","2026-08-19T02:08:40.142862Z",true,"agent",228,{"items":37},[38,43,48,53,58,63],{"id":39,"title":40,"news_slug":41,"published_at":42},"23dffa70-3b3e-452d-9730-a9c0074556ee","TMax 把「极简 RL」做成终端 Agent 工程范本:UW×Ai2 用 9B 模型跑出 27.2%,开源 14,600 训练环境","tmax-uw-ai2-terminal-agent-9b-27pct","2026-06-25T06:00:00+00:00",{"id":44,"title":45,"news_slug":46,"published_at":47},"1d80585e-c797-4aa1-ac68-ef87334d5d0c","PLaMo 3.0 Prime 正式发布：PFN 把「日语实战」做成日本国产 LLM 的差异化战场","plamo-3-0-prime-pfn-japanese-domestic","2026-06-24T08:15:00+00:00",{"id":49,"title":50,"news_slug":51,"published_at":52},"0d8fdf45-4585-47c0-9e78-3652e318b156","Apple Intelligence 中国版落地:通义千问接管语言 AI,百度负责视觉搜索","apple-intelligence-china-qwen-baidu-2026","2026-08-25T12:00:00+00:00",{"id":54,"title":55,"news_slug":56,"published_at":57},"1311adb6-dc19-41a7-a188-6760d9e53672","HF Summer 2026 报告:13 个下载量 Top 25 模型是 2022 年的老面孔","hugging-face-summer-2026-attention-adoption","2026-08-24T08:00:00+00:00",{"id":59,"title":60,"news_slug":61,"published_at":62},"1844afb1-3a1c-4acd-9e4c-f5e2792a2018","下载免费不等于商用免费：HF Summer 2026 隐藏的开源前沿许可证分水岭","frontier-license-shift-hf-summer-2026","2026-08-23T12:30:00+00:00",{"id":64,"title":65,"news_slug":66,"published_at":67},"9389d1ed-dd2d-41cb-bbc5-9a543e2b2f71","开源报告里的「参数天花板」分水岭:中国实验室把上限拉到2.78T,美国还在130B徘徊","hf-summer-2026-china-open-weight-parameter-ceiling","2026-08-20T06:00:00+00:00"]