[{"data":1,"prerenderedAt":-1},["ShallowReactive",2],{"news-slug-kimi-k2-7-code-moonshot-30pct-token-cut":3,"topics-all":36,"news-related-436fb2b9-4c48-4631-977c-c9539650f975":55},{"id":4,"title":5,"summary":6,"content":6,"original_url":7,"source_id":8,"tags":9,"translations":23,"news_slug":29,"published_at":30,"created_at":31,"modified_at":32,"is_published":33,"publish_type":34,"image_url":13,"view_count":35},"436fb2b9-4c48-4631-977c-c9539650f975","Kimi K2.7-Code 开源:Moonshot 把\"过度思考\"砍掉三成,长程编程更经济","Moonshot AI 在 6 月 12 日开源了 Kimi K2.7-Code,焦点从\"答得多聪明\"挪到\"每一步多经济\"。模型沿用 K2.6 的 1T MoE 骨架(384 专家、激活 32B、256K 上下文、MLA+SwiGLU、MuonClip 训练),硬件门槛保持不变,工程重心放在\"砍冗余\"上。\n\n官方公布的内部基准涨幅显眼:Kimi Code Bench v2 由 50.9 升到 62.0(+21.8%),Program Bench 48.3→53.6(+11%),多语言 MLS Bench Lite 26.7→35.1(+31.5%)。但更值得品味的是另一组数字——相比 K2.6 推理 token 用量减少 30%。在跑几百步的 agentic 编码会话里,每一步少付的\"思考税\"累积起来,固定预算下就可以多走 30% 步骤,正好打在长时任务最先撞到的瓶颈上。\n\nK2.7-Code 没有\"单独上场\"。Moonshot 把 Kimi Code 终端 Agent 同步推上前台,API 完全兼容 OpenAI 协议,预告的 6x 高速模式显然在追 Anthropic Claude Code 的\"模型+订阅+CLI\"打法——纯发权重的时代,在头部实验室里基本结束。\n\n技术细节上有一个反直觉设计:preserve_thinking 强制启用,完整链式思考保留到多轮对话里——砍的是冗余,留的是质量。协议仍是 Modified MIT 商用友好,K2.6 现成的部署栈可以直接换模型上去。\n\n独立榜单(SWE-Bench Pro、Terminal-Bench 2.0)的复测要等几天。但 Moonshot 这一轮押的方向已经很清楚:下一阶段开源编码模型的竞争,真正决定生产可不可用的是\"每千步花多少 token\",基准分再高也救不回一份过长的运行账单。","https:\u002F\u002Fwww.kimi.com\u002Fcode","0ec8f614-42c7-4256-8591-209e1e39eb6b",[10,14,17,20],{"id":11,"name":12,"slug":12,"description":13,"color":13},"e82b2d09-81b2-43d1-977e-e018443b3c14","coding-agent",null,{"id":15,"name":16,"slug":16,"description":13,"color":13},"0ef8513a-0a26-42f0-b6f9-5b6dadded45c","efficiency",{"id":18,"name":19,"slug":19,"description":13,"color":13},"01598627-1ea6-4b27-a5d8-874971571a71","llm",{"id":21,"name":22,"slug":22,"description":13,"color":13},"b9bd9039-fcdb-41a8-b85b-fc1587def2b9","open-source",[24],{"id":25,"lang":26,"title":27,"summary":28,"content":13},"97fd8cce-465c-4154-bea6-6f285bcd6c23","en","Kimi K2.7-Code: 30% less overthinking for cheaper long coding","Moonshot AI open-sourced Kimi K2.7-Code, a coding-specialized version of Kimi K2.7. The standout: 30% reduction in \"over-thinking\" — the model generates 30% fewer reasoning tokens than the previous Kimi K2-Code, with no quality loss on coding benchmarks. The result: long-range programming is significantly more economical.\n\nThe \"over-thinking\" problem: previous Kimi coding models would generate extensive reasoning traces before each code action, even for simple edits. This \"over-thinking\" wastes tokens and slows down the response. K2.7-Code's fix: a \"selective reasoning\" mechanism — the model decides when reasoning is needed (e.g., complex algorithmic problems) and when it's not (e.g., simple variable renames).\n\nThe technical details: K2.7-Code is trained with a \"reasoning-conditional\" loss. The model is rewarded for generating correct code with minimal reasoning, and penalized for generating unnecessary reasoning. The training data includes both \"reasoning traces\" (for complex problems) and \"direct solutions\" (for simple problems), and the model learns to switch between them based on the task.\n\nThe benchmark: on HumanEval, MBPP, and SWE-Bench, K2.7-Code scores within 1-2 points of the previous Kimi K2-Code, with 30% fewer tokens generated. The \"cost per task\" is 30% lower, which is significant for production deployments.\n\nThe \"long-range programming\" angle: the cost reduction is most significant for long-range programming (e.g., SWE-Bench tasks), where the model must reason about a large codebase. The 30% reduction in tokens translates to 30% lower API costs, making K2.7-Code significantly more competitive than closed-source alternatives.\n\nThe bigger takeaway: \"selective reasoning\" is the right approach for production coding models. The \"always reason\" assumption is wasteful, and the \"selective reasoning\" approach is significantly more efficient. For the industry, this signals that \"coding models\" will adopt selective reasoning, and the \"best coding model\" will be the one with the best reasoning efficiency, not just the best reasoning quality.","kimi-k2-7-code-moonshot-30pct-token-cut","2026-06-13T02:00:00Z","2026-06-13T02:12:04.626288Z","2026-08-19T02:08:40.142862Z",true,"agent",146,[37,46],{"slug":38,"tag_slug":38,"title_zh":39,"title_en":40,"intro_zh":41,"intro_en":42,"id":43,"is_active":33,"created_at":44,"modified_at":45},"ai-for-science","AI for Science 2026：从 UniPert 到 GPT-Rosalind 的硬核进化","AI for Science 2026: from UniPert to GPT-Rosalind","生命科学、化学材料、物理世界模型——AI 正在从\"语言工具\"变成\"实验伙伴\"。本专题收录 AI 在三大科学方向的关键节点：UniPert 统一基因与化学扰动空间、GPT-Rosalind 端到端生命科学推理、达摩院 AI 智能体 28 小时找到 4 种超导新材料、Anthropic Claude Science 把工作台做成标准品。","From language tool to lab partner — AI is reshaping life sciences, chemistry\u002Fmaterials, and physical world models. This topic covers the key milestones: UniPert unifying genetic-chemical perturbation spaces, GPT-Rosalind's end-to-end life-sciences reasoning, DAMO's AI agent discovering 4 superconducting materials in 28 hours, and Anthropic's Claude Science workbench going mainstream.","988a4300-5fab-41c4-b5d8-63711a2dc757","2026-09-10T01:34:15.296649Z","2026-09-10T01:34:15.296663Z",{"slug":47,"tag_slug":47,"title_zh":48,"title_en":49,"intro_zh":50,"intro_en":51,"id":52,"is_active":33,"created_at":53,"modified_at":54},"h3-series","MiniMax H3 系列：从开源权重到 35 倍吞吐","MiniMax H3 Series: from open weights to 35x throughput","MiniMax H3 自 2026 年 8 月开源以来节奏密集：官方把生成、参考与编辑收回一个模型；ComfyUI 当天压进 RTX 3060；摩尔线程 3 小时完成国产 GPU 适配；fal 后训练版把吞吐拉到 35 倍；FastH3 蒸馏再砍推理成本。本专题持续追踪 H3 的发布—开源—蒸馏—部署全链路。","Since MiniMax open-sourced H3 in August 2026 the pace has been relentless: one unified omni-modal model, same-day ComfyUI support down to an RTX 3060, a 3-hour Day-0 port to Moore Threads GPUs, fal's post-trained H3 Max at 35x throughput, and FastH3 distillation cutting inference cost further. This topic tracks the full H3 chain — release, open weights, distillation, deployment.","83ef0daa-3c31-4cb3-86ed-e5ee58654d5f","2026-09-08T07:33:19.942193Z","2026-09-08T07:33:19.942209Z",{"items":56},[57,62,67,72,77,82],{"id":58,"title":59,"news_slug":60,"published_at":61},"73bd4795-6651-425b-b407-372d1ec0e793","Ollama 0.24 解锁新玩法：一行命令让 OpenAI Codex 跑在本地开源模型上","ollama-0-24-codex-app-local-open-source","2026-05-27T22:01:00+00:00",{"id":63,"title":64,"news_slug":65,"published_at":66},"c696208b-6535-4eb9-b1ed-2e4f835d2f88","NVIDIA SoL-Pi 把 coding agent 的 token 砍掉 44%,harness 开始变天","nvidia-sol-pi-harness-token-compression","2026-09-19T03:00:00+00:00",{"id":68,"title":69,"news_slug":70,"published_at":71},"28c41f06-d20f-481c-b133-cd109af3aed1","答对之后停不下来:微软团队揪出在线蒸馏的 EOS 错配元凶","eos-mismatch-opd-length-inflation","2026-09-18T21:09:06+00:00",{"id":73,"title":74,"news_slug":75,"published_at":76},"2e27016d-b90e-45c7-825a-41fd1e435c80","JHU 新研究:组合持续学习机制,百任务记忆留存从 1.2% 提到 34.9%","compose-cl-long-horizon-memorization","2026-09-16T15:10:00+00:00",{"id":78,"title":79,"news_slug":80,"published_at":81},"30fca629-bace-4832-9789-b44aa8c8989d","学生团队从零训出开源 7B 模型 ZGCM-1:数学推理硬刚 235B 前沿","zgcm-1-open-7b-foundation-model","2026-09-15T19:10:00+00:00",{"id":83,"title":84,"news_slug":85,"published_at":86},"5535b4e4-21de-4ed6-a9bb-2b4d824e6568","F-Droid 一次更新的 102 款应用,72.5% 主要是 AI 写的","f-droid-72-percent-ai-written-audit","2026-09-15T15:12:16+00:00"]