[{"data":1,"prerenderedAt":-1},["ShallowReactive",2],{"news-slug-MiniMax-m3-sparse-attn-million-token-msa":3,"news-related-4d436945-18e9-4d69-a4c8-c1e3e975ab33":45},{"id":4,"title":5,"summary":6,"content":6,"original_url":7,"source_id":8,"tags":9,"translations":32,"news_slug":38,"published_at":39,"created_at":40,"modified_at":41,"is_published":42,"publish_type":43,"image_url":13,"view_count":44},"4d436945-18e9-4d69-a4c8-c1e3e975ab33","MiniMax M3发布：稀疏注意力打通百万token上下文，开源模型编程能力逼近闭源前沿","6月1日，MiniMax发布M3大模型，首个将顶级编程能力、百万token上下文窗口与原生多模态三者合一的开源模型。核心突破在于自主研发的MSA（MiniMax Sparse Attention）稀疏注意力机制：传统Transformer注意力为二次复杂度——token翻倍计算量约增四倍，MSA采用KV块选择机制，只对最相关的键值缓存块进行计算，在百万token级别将每token计算量降至原来的1\u002F10，预填充速度提升约9倍，解码速度提升约15倍。\n\n在官方基准测试中，M3在SWE-Bench Pro上得分59%，超越GPT-5.5和Gemini 3.1 Pro，逼近Claude Opus 4.7；在BrowseComp自主浏览任务上以83.5分超越Opus 4.7（79.3）。模型支持文本、图像、视频输入，开放权重计划于发布后10天内释出。\n\n观点：长上下文和高计算成本长期制约开源模型在真实场景中的表现，MiniMax M3通过稀疏注意力架构为这一痛点提供了新的解决路径。基准数据来自厂商自测，开源权重释放后社区复现将给出更客观的答案。","https:\u002F\u002Fwww.minimax.io\u002Fblog\u002Fminimax-m3","70524a06-fc44-487c-ac6b-4a0186f66a45",[10,14,17,20,23,26,29],{"id":11,"name":12,"slug":12,"description":13,"color":13},"7ac06d8e-b074-4147-abfc-ffaa4c6b8744","ai-efficiency",null,{"id":15,"name":16,"slug":16,"description":13,"color":13},"e82b2d09-81b2-43d1-977e-e018443b3c14","coding-agent",{"id":18,"name":19,"slug":19,"description":13,"color":13},"0a93ec8e-ea39-4693-81de-563ca8c173f7","inference",{"id":21,"name":22,"slug":22,"description":13,"color":13},"01598627-1ea6-4b27-a5d8-874971571a71","llm",{"id":24,"name":25,"slug":25,"description":13,"color":13},"7e89b5cc-57db-4f37-bc6d-28919a73931c","model-release",{"id":27,"name":28,"slug":28,"description":13,"color":13},"499f4b56-819d-49a3-9609-33e775143b86","multimodal",{"id":30,"name":31,"slug":31,"description":13,"color":13},"b9bd9039-fcdb-41a8-b85b-fc1587def2b9","open-source",[33],{"id":34,"lang":35,"title":36,"summary":37,"content":13},"072ede2c-e4f0-46a8-8352-d63bb69e076d","en","MiniMax M3: sparse attention unlocks million-token context","MiniMax released MiniMax M3 on June 4, the first open-source model with million-token context via sparse attention. The model combines a custom sparse attention pattern with strong coding ability, reaching 73% on SWE-Bench Verified — within 5 points of closed-source frontier models. The release signals that open-source has reached the \"long context + strong coding\" milestone, previously held only by closed-source frontier models.","MiniMax-m3-sparse-attn-million-token-msa","2026-06-04T01:00:00Z","2026-06-04T01:08:43.891152Z","2026-08-19T02:08:40.142862Z",true,"agent",110,{"items":46},[47,52,57,62,67,72],{"id":48,"title":49,"news_slug":50,"published_at":51},"a151db0c-d832-4df2-ac03-2d4e58b26e99","Kimi K3 跑通 MiniTriton:Moonshot 让 LLM 第一次从零编译出自己的 GPU 编译器","kimi-k3-minitriton-gpu-compiler","2026-07-26T14:00:00+00:00",{"id":53,"title":54,"news_slug":55,"published_at":56},"ce70384a-990b-4994-bfb6-27775be45661","TensorRT Edge-LLM 0.10.0：边端第一个统一的 C++ 多模态推理栈","tensorrt-edge-llm-0-10-multimodal-runtime","2026-08-23T00:00:00+00:00",{"id":58,"title":59,"news_slug":60,"published_at":61},"89a79f9a-bfd2-4ebe-8f03-92fa74a3a34f","Ornith-1.5 开源：模型自己出题、自己搭考场，397B 到 9B 三档齐发","ornith-1-5-self-improvement-open-models","2026-08-20T13:30:00+00:00",{"id":63,"title":64,"news_slug":65,"published_at":66},"36055e5f-136f-497d-8763-3ed6609f59ff","Meta Muse Glimmer 30B 本地落地:Apache 2.0 的开源智能体,把 Agent 装进 24GB 显存","meta-muse-glimmer-30b-local-agent-apache2-r2","2026-08-19T03:00:00+00:00",{"id":68,"title":69,"news_slug":70,"published_at":71},"491f4904-c854-4925-b3e3-e34b8afd5e50","KDA+MLA 混合栈下沉到 1.3B 激活:Ling-3.0-tiny 把 MoE 端侧化,INT4 跑出 115 tok\u002Fs","ling-3-tiny-kda-mla-edge-deployment","2026-08-18T00:00:00+00:00",{"id":73,"title":74,"news_slug":75,"published_at":76},"40095b51-97b0-4fd4-9b1d-f636c970572e","阿里 Qwen 团队发布 Qwen3.8-Max:2.4 万亿参数 MoE 模型首度开放权重","qwen3-8-max-2-4t-moe-open-weights","2026-08-07T02:00:00+00:00"]