[{"data":1,"prerenderedAt":-1},["ShallowReactive",2],{"news-slug-kimi-k2-6-programming-challenge-22-pt-beats-gpt-5-5":3,"topics-all":36,"news-related-aca19c8b-6ba8-49d8-b761-fb812fb18a77":55},{"id":4,"title":5,"summary":6,"content":6,"original_url":7,"source_id":8,"tags":9,"translations":23,"news_slug":29,"published_at":30,"created_at":31,"modified_at":32,"is_published":33,"publish_type":34,"image_url":13,"view_count":35},"aca19c8b-6ba8-49d8-b761-fb812fb18a77","Kimi K2.6 在AI编程挑战赛中夺冠：开源模型展现长时推理优势","在AI编程挑战赛「宝石单词拼图」项目中，Moonshot AI 的开源模型 Kimi K2.6 以 7 胜 1 负、22 积分横扫 GPT-5.5、Claude Opus 4.7 和 Gemini Pro 3.1 等闭源旗舰模型夺冠。\n\n比赛要求模型在 10×10 至 30×30 的滑动拼图上 10 秒内构建有效单词，长词得分、短词扣分。Kimi K2.6 以贪心策略持续滑动，在大网格中累积全场最高分。Claude Opus 4.7 在大网格上未能完成必要滑动，Muse Spark 误解扣分规则导致得分低至 -15309。\n\n完整排名：Kimi K2.6（22分）、小米 MiMo V2-Pro（20分）、GPT-5.5（16分）、GLM 5.1（15分）、Claude Opus 4.7（12分）、Gemini Pro 3.1（9分）、Grok Expert 4.2（9分）、DeepSeek V4（3分）、Muse Spark（0分）。\n\n这场比赛揭示了模型评估中被忽视的维度：长时推理的策略执行与规则理解能力，而非 benchmark 分数。开源模型正重新定义 AI 编程助手竞争格局。","https:\u002F\u002Fthinkpol.ca\u002F2026\u002F04\u002F30\u002Fan-open-weights-chinese-model-just-beat-claude-gpt-5-5-and-gemini-in-a-programming-challenge\u002F","2057406c-0696-4018-b5da-6ae1cf8ed34a",[10,14,17,20],{"id":11,"name":12,"slug":12,"description":13,"color":13},"5e628969-6d2a-437f-998a-104e4b16cfb1","ai-progress",null,{"id":15,"name":16,"slug":16,"description":13,"color":13},"e82b2d09-81b2-43d1-977e-e018443b3c14","coding-agent",{"id":18,"name":19,"slug":19,"description":13,"color":13},"0ef8513a-0a26-42f0-b6f9-5b6dadded45c","efficiency",{"id":21,"name":22,"slug":22,"description":13,"color":13},"01598627-1ea6-4b27-a5d8-874971571a71","llm",[24],{"id":25,"lang":26,"title":27,"summary":28,"content":13},"387662f6-7918-4539-95dc-09656338cb69","en","Kimi K2.6 wins the AI coding challenge on long-horizon reasoning","In the AI programming challenge \"Gemstone Word Puzzle,\" Moonshot AI's open-source model Kimi K2.6 swept closed-source flagships like GPT-5.5, Claude Opus 4.7, and Gemini Pro 3.1 to win, with 7 wins and 1 loss, 22 points.\n\nThe challenge required models to construct valid words on sliding puzzles from 10×10 to 30×30 in 10 seconds, with long words scoring positively, short words scoring negatively. Kimi K2.6 used a greedy strategy to continuously slide, accumulating the highest score on large grids. Claude Opus 4.7 failed to complete necessary slides on large grids, while Muse Spark misunderstood the scoring rules leading to a score as low as -15309.\n\nFull ranking: Kimi K2.6 (22 points), Xiaomi MiMo V2-Pro (20 points), GPT-5.5 (16 points), GLM 5.1 (15 points), Claude Opus 4.7 (12 points), Gemini Pro 3.1 (9 points), Grok Expert 4.2 (9 points), DeepSeek V4 (3 points), Muse Spark (0 points).\n\nThis challenge reveals an overlooked dimension in model evaluation: long-horizon reasoning strategy execution and rule understanding ability, rather than benchmark scores. Open-source models are redefining the AI programming assistant competitive landscape.","kimi-k2-6-programming-challenge-22-pt-beats-gpt-5-5","2026-05-04T01:01:00Z","2026-05-04T01:09:31.184760Z","2026-08-19T02:08:40.142862Z",true,"agent",154,[37,46],{"slug":38,"tag_slug":38,"title_zh":39,"title_en":40,"intro_zh":41,"intro_en":42,"id":43,"is_active":33,"created_at":44,"modified_at":45},"ai-for-science","AI for Science 2026：从 UniPert 到 GPT-Rosalind 的硬核进化","AI for Science 2026: from UniPert to GPT-Rosalind","生命科学、化学材料、物理世界模型——AI 正在从\"语言工具\"变成\"实验伙伴\"。本专题收录 AI 在三大科学方向的关键节点：UniPert 统一基因与化学扰动空间、GPT-Rosalind 端到端生命科学推理、达摩院 AI 智能体 28 小时找到 4 种超导新材料、Anthropic Claude Science 把工作台做成标准品。","From language tool to lab partner — AI is reshaping life sciences, chemistry\u002Fmaterials, and physical world models. This topic covers the key milestones: UniPert unifying genetic-chemical perturbation spaces, GPT-Rosalind's end-to-end life-sciences reasoning, DAMO's AI agent discovering 4 superconducting materials in 28 hours, and Anthropic's Claude Science workbench going mainstream.","988a4300-5fab-41c4-b5d8-63711a2dc757","2026-09-10T01:34:15.296649Z","2026-09-10T01:34:15.296663Z",{"slug":47,"tag_slug":47,"title_zh":48,"title_en":49,"intro_zh":50,"intro_en":51,"id":52,"is_active":33,"created_at":53,"modified_at":54},"h3-series","MiniMax H3 系列：从开源权重到 35 倍吞吐","MiniMax H3 Series: from open weights to 35x throughput","MiniMax H3 自 2026 年 8 月开源以来节奏密集：官方把生成、参考与编辑收回一个模型；ComfyUI 当天压进 RTX 3060；摩尔线程 3 小时完成国产 GPU 适配；fal 后训练版把吞吐拉到 35 倍；FastH3 蒸馏再砍推理成本。本专题持续追踪 H3 的发布—开源—蒸馏—部署全链路。","Since MiniMax open-sourced H3 in August 2026 the pace has been relentless: one unified omni-modal model, same-day ComfyUI support down to an RTX 3060, a 3-hour Day-0 port to Moore Threads GPUs, fal's post-trained H3 Max at 35x throughput, and FastH3 distillation cutting inference cost further. This topic tracks the full H3 chain — release, open weights, distillation, deployment.","83ef0daa-3c31-4cb3-86ed-e5ee58654d5f","2026-09-08T07:33:19.942193Z","2026-09-08T07:33:19.942209Z",{"items":56},[57,62,67,72,77,82],{"id":58,"title":59,"news_slug":60,"published_at":61},"b0c434a5-4911-4297-b1ef-2c44cbc26653","蚂蚁新研究:19769 个代码仓库,炼出百万条 agent 技能","code2skill-agent-skill-synthesis","2026-09-21T19:06:31+00:00",{"id":63,"title":64,"news_slug":65,"published_at":66},"b0c4e8d2-5662-4e3e-b489-6202eabbe97b","Dream-RSI 把历史当模拟器:162 倍杠杆重写 RSI 算力账本","dream-rsi-replay-simulator-162x","2026-09-16T06:00:00+00:00",{"id":68,"title":69,"news_slug":70,"published_at":71},"2266cea6-06f1-4932-8905-1bc3f2e5a8c0","Meta FAIR 字节蒸馏研究:End-Of-Token 渐近反超 token 蒸馏 4%,数据只需 1\u002F6","meta-fair-byte-distillation-token-ceiling-2026-09","2026-09-15T02:00:00+00:00",{"id":73,"title":74,"news_slug":75,"published_at":76},"21c7dec1-f68e-4641-974c-ae2bce87393e","教师打分、验证器掌舵:腾讯混元 FlowBalance 给自蒸馏装上方向门控,Qwen3-8B 数学均值超 GRPO 2.12 分","flowbalance-verifier-gated-self-distillation","2026-09-08T15:08:17+00:00",{"id":78,"title":79,"news_slug":80,"published_at":81},"deac2d55-76a6-40d2-8ef7-36aed2ad0105","Linux 7.2 把 AI 拉进内核开发:Sashiko 让补丁数量翻倍,Torvalds 接受「新常态」","linux-7-2-sashiko-ai-kernel-review","2026-08-20T12:00:00+00:00",{"id":83,"title":84,"news_slug":85,"published_at":86},"4bbc55d2-cabc-477f-a3ad-4e2c119aff2a","TokTier 抓住 Agent 推理的隐藏瓶颈：缓存命中 94.1%，分词仍吃掉 64% 首 token 时间","toktier-stateful-tokenization-agent-serving","2026-07-31T17:56:30+00:00"]