[{"data":1,"prerenderedAt":-1},["ShallowReactive",2],{"news-slug-meta-muse-glimmer-30b-local-agent-apache2-r2":3,"news-related-36055e5f-136f-497d-8763-3ed6609f59ff":41},{"id":4,"title":5,"summary":6,"content":7,"original_url":8,"source_id":9,"tags":10,"translations":27,"news_slug":34,"published_at":35,"created_at":36,"modified_at":37,"is_published":38,"publish_type":39,"image_url":14,"view_count":40},"36055e5f-136f-497d-8763-3ed6609f59ff","Meta Muse Glimmer 30B 本地落地:Apache 2.0 的开源智能体,把 Agent 装进 24GB 显存","Meta Superintelligence Labs 8 月 10 日在 Apache 2.0 下开源 Muse Glimmer 30B,主打本地常驻智能体场景;4-bit 量化后模型不到 20 GB,可在 RTX 5090 或 M4\u002FM5 Max Mac 上跑出实时的 agent 交互。","当模型从云端搬到一台笔记本,「本地能跑」和「本地值得跑」之间那条线,被 Meta 在 8 月 10 日划过了。Meta Superintelligence Labs 在 Apache 2.0 许可下开源了 Muse Glimmer 30B,这是一个参数量约 296 亿的稠密多模态模型,从更大的教师模型 Muse Spark 蒸馏而来,专门优化「常驻本地的智能体工作流」——不是聊天、不是角色扮演,是端到端任务执行、工具调用、长链路规划与失败恢复。\n\n最值得关注的不是 30B 这个数字,而是许可证。Meta 从 Llama 时代起一直用自家「Llama Community Licence」,附有可接受使用政策、月活用户阈值与品牌条款;Glimmer 直接换成标准 OSI 认可的开源许可,商用、微调、蒸馏、再分发均无需申请。对于在企业法务走完 Llama 合规审查前一直被卡住的团队来说,这才是这次发布真正改变的东西。\n\n模型架构是 Meta 有意选择的「反潮流」:稠密、不是 MoE。52 层解码器、3 层滑动窗口 + 1 层全局注意力重复 13 次、gated grouped-query attention(GQA ratio 16:1),独立的 ViT-G\u002F14 感知编码器约 18 亿参数。显存配置上,BF16 全精度需要 55 GB 以上;Meta 自带的 K-Quant 校准 4-bit 把语言模型压到 20 GB 以内,给 KV cache、视觉编码器、speculative-decoding drafter 留出空间,刚好塞进 24 GB 或 32 GB 的消费级显存预算。Meta 报告的 4-bit 平均 benchmark 退化只有 0.2–1.0%。\n\n推理速度靠打包的 DFlash 投机解码 drafter 实现。Meta 给出的数据:RTX 5090 上从 74.9 tok\u002Fs 拉到 233.4 tok\u002Fs(3.1×),M5 Max 从 26.6 拉到 50.2(1.8×),M4 Max 从 23.7 拉到 37.8(1.5×)。这意味着在不联网、不买云 token 的前提下,一台 5090 工作站或 M5 Max MacBook 就能跑出「实时的 agent 交互」体感。\n\n基准对比选得很诚实:Glimmer 对标 Gemma4-31B 和 Qwen3.6-27B,不与闭源前沿模型直接打。在 MCP Atlas(75.5 vs Gemma4 的 54.2)、WildClawBench(47.6 vs 37.6)、AA-LCR 长上下文(80.0 vs 68.3)上拉开明显差距——这些都是 agentic、工具编排、长链路场景;在 SWE-Bench Pro 上 51.2,优于 Gemma4 的 36.9,与 Qwen3.6 的 50.2 几乎打平。在 AIME 2026 数学上 94.7,在 ChartXiv 上 78.8——这三类任务与 Qwen 的差距都收得很紧,意味着本地模型真正进入「可以做严肃工作」的区间,但仍未到能完全替代前沿闭源模型的水平。\n\n生态落地节奏也很快:发布当天 Transformers 即可用,llama.cpp、MLX、ExecuTorch、vLLM(走 transformers 后端)、SGLang、Ollama、LM Studio、Unsloth 陆续跟进,Together AI \u002F Fireworks AI \u002F OpenRouter 作为托管伙伴上线;AMD、Arm、Dell、Intel、NVIDIA 都已在做硬件适配。\n\n所以呢?Glimmer 不会替代 Claude 或 GPT 做最难的推理,但它把「本地 agent 模型终于可用」这件事坐实了:在 24–32 GB 的显存预算下,一个能稳定调工具、看截图、做多步规划的模型第一次真正装进了笔记本。配合 Apache 2.0,私有部署、air-gap 场景、企业微调的法律摩擦被显著降低,做本地优先架构或者把 agent 路由拆成「本地高频 + 云端难例」的团队,值得把它放进评估集。","https:\u002F\u002Fwww.mindstudio.ai\u002Fblog\u002Fmeta-muse-glimmer-30b-open-weights","a1f0bda7-5035-4317-b63b-72693539d2e3",[11,15,18,21,24],{"id":12,"name":13,"slug":13,"description":14,"color":14},"6ad31a14-c0da-42df-81fd-564281f768db","agentic-ai",null,{"id":16,"name":17,"slug":17,"description":14,"color":14},"0a93ec8e-ea39-4693-81de-563ca8c173f7","inference",{"id":19,"name":20,"slug":20,"description":14,"color":14},"01598627-1ea6-4b27-a5d8-874971571a71","llm",{"id":22,"name":23,"slug":23,"description":14,"color":14},"7e89b5cc-57db-4f37-bc6d-28919a73931c","model-release",{"id":25,"name":26,"slug":26,"description":14,"color":14},"b9bd9039-fcdb-41a8-b85b-fc1587def2b9","open-source",[28],{"id":29,"lang":30,"title":31,"summary":32,"content":33},"f37cb413-af85-4cdd-979d-154aae19b4d0","en","Meta's Muse Glimmer 30B lands locally: Apache 2.0 open-weights agent on a single 24GB GPU","Meta Superintelligence Labs open-sourced Muse Glimmer 30B under Apache 2.0 on August 10, built for always-on local agent workflows. The 4-bit artifact lands under 20 GB and delivers fluid agent interaction on an RTX 5090 or M4\u002FM5 Max MacBook.","Meta drew the line between \"runs locally\" and \"is worth running locally\" on August 10. Meta Superintelligence Labs open-sourced Muse Glimmer 30B under Apache 2.0, a ~29.6B dense multimodal model distilled from the larger Muse Spark teacher, optimized specifically for always-on local agent workflows — not chat, not roleplay, but end-to-end task execution, tool use, long-horizon planning, and failure recovery.\n\nThe headline is not the 30B parameter count. It is the license. Meta has shipped open weights under its own Llama Community Licence for years, complete with an acceptable-use policy, monthly-active-user threshold, and branding clause. Glimmer switches to a standard OSI-approved permissive license: commercial use, fine-tuning, distillation, and redistribution all permitted without asking. For teams that have been blocked behind a Llama licence review, that is the change that actually matters.\n\nThe architecture is a deliberate counter to the prevailing fashion: dense, not MoE. A 52-layer decoder with three sliding-window layers + one global-attention layer repeating thirteen times, gated grouped-query attention at a 16:1 GQA ratio, and an independent ~1.8B-parameter ViT-G\u002F14 perception encoder. On memory, BF16 full precision needs 55 GB or more; Meta's K-Quant calibrated 4-bit pushes the language model under 20 GB, leaving headroom for the KV cache, the vision encoder, and the speculative-decoding drafter inside a 24 GB or 32 GB consumer envelope. Meta reports 4-bit average benchmark degradation of only 0.2–1.0%.\n\nInference speed is delivered through a bundled DFlash speculative-decoding drafter. Meta's numbers: 74.9 → 233.4 tok\u002Fs on RTX 5090 (3.1×), 26.6 → 50.2 on M5 Max (1.8×), 23.7 → 37.8 on M4 Max (1.5×). The practical upshot: a 5090 workstation or an M5 Max MacBook can deliver fluid, real-time agent interaction without any cloud token spend or network call.\n\nThe benchmark comparison is framed honestly. Glimmer is pitted against Gemma4-31B and Qwen3.6-27B, not against closed-source frontier models. On MCP Atlas (75.5 vs Gemma4's 54.2), WildClawBench (47.6 vs 37.6), and AA-LCR long-context (80.0 vs 68.3), it pulls a meaningful lead — those are agentic, tool-orchestration, and long-horizon tasks. On SWE-Bench Pro it scores 51.2, well ahead of Gemma4's 36.9 and effectively tied with Qwen3.6's 50.2. On AIME 2026 math it scores 94.7, on ChartXiv 78.8 — both close to Qwen, narrow enough that workload decides the winner, not the table. Net read: local models have entered the \"can do serious work\" band; they have not entered the \"can replace the frontier\" band.\n\nEcosystem rollout is fast. Transformers works day one; llama.cpp, MLX, ExecuTorch, vLLM (via the transformers backend), SGLang, Ollama, LM Studio, and Unsloth land in the following days; Together AI, Fireworks AI, and OpenRouter are named as hosted launch partners; AMD, Arm, Dell, Intel, and NVIDIA are working on hardware optimization.\n\nSo what? Glimmer does not displace Claude or GPT on the hardest reasoning work. What it does is make \"local agent model finally usable\" real: for the first time, a model stable enough at tool calling, screenshot reading, and multi-step planning actually fits inside a laptop's 24–32 GB memory budget. Combined with Apache 2.0, the legal friction around private deployment, air-gapped environments, and enterprise fine-tuning drops sharply. Teams building local-first architectures, or splitting agent routing into \"local high-volume + cloud hard cases\", should put Glimmer on the eval bench.","meta-muse-glimmer-30b-local-agent-apache2-r2","2026-08-19T03:00:00Z","2026-08-19T11:07:18.022643Z","2026-08-19T11:07:18.022652Z",true,"agent",71,{"items":42},[43,48,53,58,63,68],{"id":44,"title":45,"news_slug":46,"published_at":47},"f26ace13-9c96-47ea-a528-b6682a22aa1e","Apodex 1.1 把推理搬进真实执行:PIVOT-RL 定位关键决策点,35B mini 开源","apodex-1-1-agentic-execution-pivot-rl","2026-08-25T14:30:00+00:00",{"id":49,"title":50,"news_slug":51,"published_at":52},"b4754043-6b19-499f-8459-f8fc786f4d80","Pokee-Isaac 28B 把 10M 上下文塞进客户边界:28B 参数在 RULER 10M 上 93.3%","pokee-isaac-28b-10m-context","2026-08-20T14:00:00+00:00",{"id":54,"title":55,"news_slug":56,"published_at":57},"5878a668-282c-4b88-b2b8-7eef40b7938c","LFM2.5-2.6B：2.5GB 内存跑本机 Agent 220 tok\u002Fs","lfm2-5-2-6b-on-device-agent","2026-08-11T00:00:00+00:00",{"id":59,"title":60,"news_slug":61,"published_at":62},"79c1684f-f61d-4799-b3d0-6450c4ad10e8","Muse Glimmer:Meta 把 30B 「常驻本地的智能体」开源,把 Agent 拉到笔记本里 7×24 跑","muse-glimmer-30b-open-agentic-local","2026-08-10T00:00:00+00:00",{"id":64,"title":65,"news_slug":66,"published_at":67},"c94766df-827e-4e4e-a006-b6639ec76722","DeepSeek V4-Flash-0731 转正观察:权重不动,后训练把 Agent 分数打到 V4-Pro 之上","deepseek-v4-flash-0731-agent-benchmark-official-aug2026","2026-08-01T02:00:00+00:00",{"id":69,"title":70,"news_slug":71,"published_at":72},"5ef8c5da-7221-4e22-bd46-c1ef55c3120d","韩国 Solar Open 2：Linear Attention 进 MoE，250B\u002F15B 激活","upstage-solar-open-2","2026-07-23T03:30:00+00:00"]