[{"data":1,"prerenderedAt":-1},["ShallowReactive",2],{"news-slug-meta-llama-4-scout-17b-109b-10m-irope":3,"news-related-39cbd58f-1761-40ff-9925-7f7b1dcf3c5d":36},{"id":4,"title":5,"summary":6,"content":6,"original_url":7,"source_id":8,"tags":9,"translations":23,"news_slug":29,"published_at":30,"created_at":31,"modified_at":32,"is_published":33,"publish_type":34,"image_url":13,"view_count":35},"39cbd58f-1761-40ff-9925-7f7b1dcf3c5d","Llama 4 Scout：Meta 首款 MoE 开源 VLM，10M 上下文重新定义边缘推理","4月中旬，Meta 正式发布 Llama 4 系列，首次将 Mixture-of-Experts（MoE）架构引入 Llama 家族。系列中定位最「轻」的 Scout（17B 激活参数 \u002F 109B 总参）已开源，引发开发者社区广泛讨论——不是因为它最大，而是因为它首次让十亿级上下文 + 视觉理解 + 单卡部署成为可能。\n\nScout 最引人注目的技术指标是 10M token 上下文窗口。传统 RoPE 在超长序列上信噪比下降明显，Meta 的解法是 iRoPE：在第 1、2、3 层使用标准 RoPE 保留局部 token 顺序；在第 4 层切换为 NoPE，移除绝对位置编码，让注意力头对整个因果掩码做全局感知。MoE 稀疏激活设计让 Scout 虽有 109B 总参，但每个 token 只需激活 17B——用 17B 算力获得近似 100B+ 模型的知识容量。\n\n基准测试显示：MMLU-Pro Maverick 80.5 分超越 GPT-4o（78.0），ChartQA\u002FDocVQA 创同规模 SOTA，Scout 在 10M token NIAH 测试维持 >99% 准确率，而竞品在 128K–1M 区间就开始「撞墙」。但在纯 STEM 推理上 OpenAI o 系列仍领先。\n\nMeta 同步开源了 Llama Guard 4（12B）和 Prompt Guard 2（86M），构成四层安全 pipeline。Llama 4 Scout 的出现把一个信号进一步强化：开源模型的竞争焦点正从「参数量」转向「效率密度」。","https:\u002F\u002Fai.meta.com\u002Fblog\u002Fllama-4\u002F","a1f0bda7-5035-4317-b63b-72693539d2e3",[10,14,17,20],{"id":11,"name":12,"slug":12,"description":13,"color":13},"0a93ec8e-ea39-4693-81de-563ca8c173f7","inference",null,{"id":15,"name":16,"slug":16,"description":13,"color":13},"01598627-1ea6-4b27-a5d8-874971571a71","llm",{"id":18,"name":19,"slug":19,"description":13,"color":13},"499f4b56-819d-49a3-9609-33e775143b86","multimodal",{"id":21,"name":22,"slug":22,"description":13,"color":13},"b9bd9039-fcdb-41a8-b85b-fc1587def2b9","open-source",[24],{"id":25,"lang":26,"title":27,"summary":28,"content":13},"c642f826-d3e8-4994-b8f4-6e76009e9036","en","Llama 4 Scout: Meta's first MoE VLM with 10M context","In mid-April, Meta officially released the Llama 4 series, bringing the Mixture-of-Experts (MoE) architecture to the Llama family for the first time. The lightest-positioned model in the series, Scout (17B activated \u002F 109B total), is now open-source, sparking wide developer community discussion — not because it's the largest, but because it makes for the first time billion-scale context + vision understanding + single-card deployment possible.\n\nScout's most striking technical spec is the 10M token context window. Traditional RoPE's signal-to-noise ratio drops sharply on ultra-long sequences; Meta's solution is iRoPE: standard RoPE on layers 1, 2, and 3 to preserve local token order; switching to NoPE on layer 4, removing absolute positional encoding, letting attention heads do global perception on the entire causal mask. The MoE sparse-activation design means that although Scout has 109B total parameters, each token only needs to activate 17B — getting near 100B+ model knowledge capacity at 17B compute.\n\nBenchmark results show: MMLU-Pro Maverick at 80.5 surpasses GPT-4o (78.0); ChartQA\u002FDocVQA set new SOTA at the same scale; Scout maintains >99% accuracy on 10M token NIAH tests, while competitors start hitting walls in the 128K-1M range. But on pure STEM reasoning, OpenAI's o-series still leads.\n\nMeta simultaneously open-sourced Llama Guard 4 (12B) and Prompt Guard 2 (86M), forming a four-layer safety pipeline. Llama 4 Scout's emergence reinforces a signal: the open-source model's competitive focus is shifting from \"parameter count\" to \"efficiency density.\"","meta-llama-4-scout-17b-109b-10m-irope","2026-04-27T04:10:00Z","2026-04-27T04:08:06.601623Z","2026-08-19T02:08:40.142862Z",true,"agent",121,{"items":37},[38,43,48,53,58,63],{"id":39,"title":40,"news_slug":41,"published_at":42},"ce70384a-990b-4994-bfb6-27775be45661","TensorRT Edge-LLM 0.10.0：边端第一个统一的 C++ 多模态推理栈","tensorrt-edge-llm-0-10-multimodal-runtime","2026-08-23T00:00:00+00:00",{"id":44,"title":45,"news_slug":46,"published_at":47},"7bdf2df4-bb0f-42ae-b0ff-187de2e3558e","把语音 AI 拆成开源乐高：HF + Cerebras 用 Gemma 4 + Qwen3-TTS 拼出实时对话流水线","hf-cerebras-voice-ai","2026-07-02T02:00:00+00:00",{"id":49,"title":50,"news_slug":51,"published_at":52},"68b6bbbc-384c-43fb-8024-2cf050107149","稀疏MoE+投机解码：开源模型首次在推理速度上超越闭源方案","stepfun-step-3-7-flash-409-tps-198b-moe","2026-06-04T13:00:00+00:00",{"id":54,"title":55,"news_slug":56,"published_at":57},"4d436945-18e9-4d69-a4c8-c1e3e975ab33","MiniMax M3发布：稀疏注意力打通百万token上下文，开源模型编程能力逼近闭源前沿","MiniMax-m3-sparse-attn-million-token-msa","2026-06-04T01:00:00+00:00",{"id":59,"title":60,"news_slug":61,"published_at":62},"ff0bc92a-295a-4707-be8d-76115fe9eeee","PerceptionBench 出炉:16 个前沿多模态模型,视觉感知无一及格","moonshot-perceptionbench-atomic-perception","2026-08-26T13:15:00+00:00",{"id":64,"title":65,"news_slug":66,"published_at":67},"cac485ab-a429-4ebc-88c6-a1f924f978ff","AWS 开源 KeysAndValues:微调时就让模型学会“遗忘”,单张 A100 撑住 128K","aws-keysvalues-sparse-attention-finetuning","2026-08-26T05:20:00+00:00"]