[{"data":1,"prerenderedAt":-1},["ShallowReactive",2],{"news-slug-xiaomi-mimo-code-compute-memory-evolution":3,"news-related-61222f5d-b7e6-4200-8517-3b1972040d24":36},{"id":4,"title":5,"summary":6,"content":6,"original_url":7,"source_id":8,"tags":9,"translations":23,"news_slug":29,"published_at":30,"created_at":31,"modified_at":32,"is_published":33,"publish_type":34,"image_url":13,"view_count":35},"61222f5d-b7e6-4200-8517-3b1972040d24","小米开源 MiMo Code：把 Coding Agent 拆成「计算-记忆-演化」三段式，长程编程首次跑通","当 Coding Agent 的执行步数从十步冲到几十甚至几百步，错误会一步步累积，而过程中又没有外部纠错信号——这正是长程编程的真正瓶颈。小米 MiMo 团队 6 月 10 日开源的 **MiMo Code**，没有在「更聪明的模型」上押注，而是把整套 harness 拆成 **计算、记忆、演化** 三段时间尺度分别优化：\n\n**第一段：计算（单轮决策质量）。** MiMo Code 引入两个正交的 test-time compute 杠杆。Max Mode 每轮并行生成 N=5 个候选解（temperature=1），再让同一模型作裁判挑最优，在 SWE-Bench Pro 上比单采样提升 10–20%，代价是 4–5 倍 token；Goal 则是独立终止校验器——用户写下「测试全过且已提交」之类的自然语言收尾条件，每当 Agent 想结束时系统自动调一次独立 model 对照上下文判定，避免自动跑里常见的「假装做完」。两者可同时开启。\n\n**第二段：记忆（任务内的状态连续性）。** 团队明确指出，靠「压缩历史」是死路——远端信息会被反复稀释，更像 Mamba 的局限而非 Transformer 的劣势。MiMo Code 改为显式存储-检索结构：什么信息值得写入持久层、何时被召回，由 harness 决定，让模型真正具备「按需回看」能力。\n\n**第三段：演化（跨 session 经验蒸馏）。** 不同任务里沉淀的失败-修复模式应当回流到 prompt 或工具策略，而不是每轮从零开始。\n\n对比同期动辄堆 GPU 集群的方案，MiMo Code 的工程哲学更贴近软件工程本身——把可靠性当 **过程** 设计，而不是依赖模型一夜变聪明。对国内 Agent 开源生态，这或许比单一 benchmark 上的 SOTA 更值得跟。\n\n（来源：小米 MiMo 官方博客《MiMo Code: Scaling Coding Agents to Long-Horizon Tasks》，2026-06-10）","https:\u002F\u002Fmimo.xiaomi.com\u002Fblog\u002Fmimo-code-long-horizon","581853c1-b1f6-420b-9124-243143660e92",[10,14,17,20],{"id":11,"name":12,"slug":12,"description":13,"color":13},"a8002d98-9df1-4ab9-94d4-a7625af634c4","china-ai",null,{"id":15,"name":16,"slug":16,"description":13,"color":13},"e82b2d09-81b2-43d1-977e-e018443b3c14","coding-agent",{"id":18,"name":19,"slug":19,"description":13,"color":13},"01598627-1ea6-4b27-a5d8-874971571a71","llm",{"id":21,"name":22,"slug":22,"description":13,"color":13},"b9bd9039-fcdb-41a8-b85b-fc1587def2b9","open-source",[24],{"id":25,"lang":26,"title":27,"summary":28,"content":13},"6d80aeaa-3647-4f31-b38d-0a68649195d0","en","Xiaomi's MiMo Code: coding agents in three evolutionary stages","When a Coding Agent's execution steps go from 10 to dozens or even hundreds, errors accumulate step by step, and there is no external error-correction signal in the process — that is the real bottleneck of long-horizon programming. Xiaomi's MiMo team's open-source MiMo Code (June 10) does not bet on \"a smarter model,\" but instead splits the entire harness into three separate time-scale optimizations — **compute, memory, evolution**:\n\n**Stage one: compute (single-turn decision quality).** MiMo Code introduces two orthogonal test-time compute levers. Max Mode generates N=5 candidate solutions in parallel per turn (temperature=1), then lets the same model serve as judge to pick the best — this gives a 10-20% boost on SWE-Bench Pro over single sampling, at the cost of 4-5× tokens. Goal is an independent termination verifier — when the user writes a natural-language closure condition like \"tests all pass and submitted,\" every time the Agent wants to end, the system automatically calls an independent model to check against context, avoiding the common \"pretend done\" in auto-runs. Both can be enabled simultaneously.\n\n**Stage two: memory (state continuity within a task).** The team explicitly points out that relying on \"compress history\" is a dead end — distant information gets diluted repeatedly, more like Mamba's limitation than Transformer's disadvantage. MiMo Code switches to an explicit store-retrieve structure: what information is worth writing to persistent storage, and when it is recalled, are decided by the harness, giving the model true \"look back on demand\" capability.\n\n**Stage three: evolution (cross-session experience distillation).** Failure-repair patterns accumulated in different tasks should flow back into the prompt or tool strategy, not start from scratch every round.\n\nCompared to same-period solutions that pile on GPU clusters, MiMo Code's engineering philosophy is closer to software engineering itself — designing reliability as a **process**, not relying on the model to get smarter overnight. For the domestic Agent open-source ecosystem, this may be more worth following than a single SOTA on a benchmark.","xiaomi-mimo-code-compute-memory-evolution","2026-06-11T04:00:00Z","2026-06-11T04:10:20.347262Z","2026-08-19T02:08:40.142862Z",true,"agent",117,{"items":37},[38,43,48,53,58,63],{"id":39,"title":40,"news_slug":41,"published_at":42},"49d0aaf8-6fcf-4bf9-83fb-19a016ae2784","CompactionRL:把上下文压缩塞进 RL 循环,GLM-5.2 训练管线吃下 5–7pp 编码代理增益","compactionrl-glm-5-2","2026-07-10T12:08:00+00:00",{"id":44,"title":45,"news_slug":46,"published_at":47},"fa8a0c53-dd89-4870-8e6e-978ce038a919","GitHub Copilot 上线 Kimi K2.7：开源权重模型首次进入默认模型选择器","github-copilot-kimi-k2-7","2026-07-03T06:00:00+00:00",{"id":49,"title":50,"news_slug":51,"published_at":52},"fc0cd2d4-cc5f-4f97-b91f-08719e41e8ec","Qwen3.6-27B：27B密集模型超越397B MoE，单卡部署的编程新选择","qwen-3-6-27b-dense-beats-397b-moe-coding-77pct","2026-04-24T03:30:00+00:00",{"id":54,"title":55,"news_slug":56,"published_at":57},"d4d40e4b-04e9-45cf-a04c-792aca45b152","Kimi K2.6开源发布：万亿参数MoE模型的长时编程与Agent Swarm突破","kimi-k2-6-trillion-moe-1t-32b-active-256k","2026-04-24T03:00:00+00:00",{"id":59,"title":60,"news_slug":61,"published_at":62},"f6e4aab0-7693-4c2c-bb66-c1641fc2cc3e","Ox Alpha 谜底揭晓:智谱 GLM-5.3-Flash,MIT 开源 320B MoE","ox-alpha-glm-5-3-flash-reveal","2026-08-27T13:30:00+00:00",{"id":64,"title":65,"news_slug":66,"published_at":67},"804ab59a-a8d6-4b61-bf74-8f6f2bdae83c","智谱把 Flash 做成一件正经事:一次说清 GLM-5.3-Flash 的架构和 benchmark 真相","glm-5-3-flash-hybrid-attention-architecture","2026-08-27T08:00:00+00:00"]