[{"data":1,"prerenderedAt":-1},["ShallowReactive",2],{"news-slug-metrollm-bench-transit-kiosk-llm":3,"topics-all":38,"news-related-2731ed1c-17c3-4d85-9174-983cf50743e3":57},{"id":4,"title":5,"summary":6,"content":7,"original_url":8,"source_id":9,"tags":10,"translations":24,"news_slug":31,"published_at":32,"created_at":33,"modified_at":34,"is_published":35,"publish_type":36,"image_url":14,"view_count":37},"2731ed1c-17c3-4d85-9174-983cf50743e3","地铁售票机上的 AI 大考:2.6GB 端侧模型 91.32 分超 GPT-5.6,规则基线也拿 84.6","Continker 开源 955 用例的地铁售票机基准:2.6GB 的 Qwen 3.5 4B 学生模型 Tier-1 拿 91.32,超 GPT-5.6 两档、打平 GPT-5.4;PEFT 增益从 2B 的 +7.03 递减到 27B 的 -0.91,规则基线也拿 84.6。","地铁售票机要不要接大模型?荷兰工作室 Continker 用一份 955 用例的基准测试给出了目前最完整的证据链。论文 9 月 9 日挂上 arXiv(编号 2609.10016),代码、数据集、微调模型全部开源,Apache 2.0 协议。\n\n## 测的是什么\n\nMetroLLM-Bench 把语言模型放进售票机的位置:读一段自然语言写成的运营策略,通过工具调用处理路线规划、票价计算、故障应对、无障碍服务和对抗输入,最后输出机器可校验的终端状态。覆盖亚特兰大、多哈、旧金山、台北、芝加哥、北京六座真实地铁系统,站点数从 37 到 414,计费模型分单一票价、里程计价、混合计价三类,共 11 个类别的用例。评分分两层:Tier 1 有 14 个确定性组件,核路线、票价、工具调用对错,不需要任何 API key;Tier 2 有 8 个语义组件,其中 6 个由 Claude Haiku 当裁判。955 个用例按 75\u002F25 分层切分,238 个用例严格留出,只用来报告成绩。26 个来自 6 家厂商的模型参赛,23 个上榜。\n\n## 核心数字\n\n留出集上,一个 2.6GB 的 Qwen 3.5 4B 学生模型(PEFT 微调,Q4_K_M 量化)拿到 Tier 1 91.32 分,超过 GPT-5.6 两档(luna 中等推理 90.63、sol 最大推理 90.00),与 GPT-5.4 全量版最大推理只差 0.05 分。更扎眼的是对照线:一个纯规则的确定性基线拿到 84.6。也就是说,大模型对规则系统的净优势只有 6.7 分左右,且集中在策略适配、复合场景、无障碍和时间推理这些\"规则写不完\"的地方。综合榜前六里只有两个 OpenAI 专有模型,榜首是 Muse Glimmer 30B 的 92.03。\n\n## PEFT 的边际递减\n\n这份测试最有价值的可能不是榜单,而是微调增益随基座变大而单调衰减的曲线:2B 基座微调后平均涨 7.03 分,4B 涨 2.00,9B 涨 1.65,到 27B 反而跌 0.91——基座越强,适配器能做的越少,甚至帮倒忙。955 全量矩阵的 bootstrap 置信区间证实了 4B 的增益(+1.72)和 27B 的回退(-1.09)都显著。复现门槛也压得很低:一台 M2 Max 跑 15 用例探针约 62 tok\u002Fs、占用 7GB 内存,README 给了从 llama.cpp 部署到全量复现的三档路径。\n\n## 两点保留\n\n一是 Hugging Face 社区里已有讨论指出:基准主要打分最终状态,工具调用失败后的恢复行为(查询超时、返回脏数据、读卡器重刷)没有单独计分层,而真实售票机恰恰死在这些地方。二是 Tier 2 的 LLM 裁判与人工标注只有 82% 完全一致,论文自己也把比较性结论压在确定性的 Tier 1 上。另外 serving 配置本身就能让 Qwen 3.5 与 3.8 的对比移动 2.7 分——部署细节的权重不亚于选型。\n\n结论可以这样收:在范围收窄、规则密集的垂直任务上,2.6GB 的本地模型已经够用,数据主权场景(公交、医疗、政务)的端侧 AI 采购可以拿这份基准当谈判起点;但\"打平旗舰\"的边界也在数字里——离开地铁这个沙盒,优势未必还在。\n\n参考:arXiv:2609.10016(https:\u002F\u002Farxiv.org\u002Fabs\u002F2609.10016);GitHub:continker\u002Fmetrollm-bench(https:\u002F\u002Fgithub.com\u002Fcontinker\u002Fmetrollm-bench)","https:\u002F\u002Farxiv.org\u002Fabs\u002F2609.10016","7437aeb9-930c-4866-a2e9-48003c1a792b",[11,15,18,21],{"id":12,"name":13,"slug":13,"description":14,"color":14},"120fa59a-ff6f-4537-9bf5-f818df636a0e","benchmark",null,{"id":16,"name":17,"slug":17,"description":14,"color":14},"0ef8513a-0a26-42f0-b6f9-5b6dadded45c","efficiency",{"id":19,"name":20,"slug":20,"description":14,"color":14},"01598627-1ea6-4b27-a5d8-874971571a71","llm",{"id":22,"name":23,"slug":23,"description":14,"color":14},"b9bd9039-fcdb-41a8-b85b-fc1587def2b9","open-source",[25],{"id":26,"lang":27,"title":28,"summary":29,"content":30},"85498be5-a1c7-4156-997c-b1d46a6ac9f8","en","MetroLLM-Bench: 2.6GB on-device model beats GPT-5.6 at kiosks","A 955-case transit-kiosk benchmark: a 2.6GB Qwen 3.5 4B student hits 91.32 Tier-1, beating both GPT-5.6 tiers; PEFT gains decay from +7.03 to -0.91.","Should a transit kiosk run a large language model? Dutch studio Continker offers the most complete evidence chain yet: a 955-case benchmark, posted to arXiv on Sep 9 (2609.10016), with code, dataset, and fine-tuned models all open-sourced under Apache 2.0.\n\n## What it tests\n\nMetroLLM-Bench puts a language model in the kiosk's seat: it reads an operations policy written in prose, then handles routing, fare calculation, disruptions, accessibility, and adversarial input through tool calls, ending each case in a machine-checkable terminal state. Six real metro systems are covered — Atlanta, Doha, San Francisco, Taipei, Chicago, Beijing — spanning 37 to 414 stations and three fare models, across eleven case categories. Scoring has two tiers: Tier 1 packs 14 deterministic components covering route, fare, and tool-call correctness, needing no API key; Tier 2 adds 8 semantic components, six judged by Claude Haiku. The 955 cases are split 75\u002F25 with 238 strictly held out for reporting. Twenty-six models from six vendors entered; twenty-three are ranked.\n\n## The headline numbers\n\nOn the held-out set, a 2.6GB Qwen 3.5 4B student (PEFT-tuned, Q4_K_M quantized) scores 91.32 on Tier 1 — above both GPT-5.6 tiers (luna at medium reasoning 90.63, sol at maximum 90.00) and within 0.05 points of GPT-5.4 full at maximum effort. The more sobering line is the control: a pure rule-based deterministic baseline reaches 84.6. The LLM's net edge over rules is roughly 6.7 points, concentrated in policy adaptation, compound scenarios, accessibility, and temporal reasoning — the parts rules cannot fully enumerate. Only two proprietary models crack the composite top six, both OpenAI; Muse Glimmer 30B leads at 92.03.\n\n## The diminishing returns of PEFT\n\nPerhaps the most valuable output is not the leaderboard but the curve: fine-tuning gains decay monotonically as the base model grows. The 2B base gains +7.03 after PEFT, 4B gains +2.00, 9B +1.65, and 27B actually drops -0.91 — the stronger the base, the less an adapter can help, and past a point it hurts. Bootstrap CIs on the full 955-case matrix certify both the 4B gain (+1.72) and the 27B regression (-1.09). Reproduction is deliberately cheap: an M2 Max runs the 15-case probe at about 62 tok\u002Fs in 7GB of RAM, and the README offers three tiers from llama.cpp serving to full replication.\n\n## Two caveats\n\nFirst, community discussion on Hugging Face already flags that the benchmark mostly scores final states; recovery behavior after tool-call failures — a timed-out fare lookup, a schedule API returning garbage, a card reader needing a second tap — has no separate scoring tier, and real kiosks die exactly there. Second, the Tier 2 LLM judge agrees with human annotators only 82% of the time, and the paper itself leans comparative claims on deterministic Tier 1. Serving configuration alone moves the Qwen 3.5-vs-3.8 comparison by 2.7 points — deployment details weigh as much as model choice.\n\nThe takeaway: on narrow, rule-dense vertical tasks, a 2.6GB local model is already good enough, and sovereign-deployment buyers (transit, healthcare, government) can treat this benchmark as a negotiating baseline. But the boundary of \"matching the frontier\" is also in the numbers — step outside the metro sandbox and the edge may not travel.\n\nRefs: arXiv:2609.10016 (https:\u002F\u002Farxiv.org\u002Fabs\u002F2609.10016); GitHub: continker\u002Fmetrollm-bench (https:\u002F\u002Fgithub.com\u002Fcontinker\u002Fmetrollm-bench)","metrollm-bench-transit-kiosk-llm","2026-09-12T23:08:18Z","2026-09-12T23:08:32.245467Z","2026-09-12T23:08:32.245475Z",true,"agent",76,[39,48],{"slug":40,"tag_slug":40,"title_zh":41,"title_en":42,"intro_zh":43,"intro_en":44,"id":45,"is_active":35,"created_at":46,"modified_at":47},"ai-for-science","AI for Science 2026：从 UniPert 到 GPT-Rosalind 的硬核进化","AI for Science 2026: from UniPert to GPT-Rosalind","生命科学、化学材料、物理世界模型——AI 正在从\"语言工具\"变成\"实验伙伴\"。本专题收录 AI 在三大科学方向的关键节点：UniPert 统一基因与化学扰动空间、GPT-Rosalind 端到端生命科学推理、达摩院 AI 智能体 28 小时找到 4 种超导新材料、Anthropic Claude Science 把工作台做成标准品。","From language tool to lab partner — AI is reshaping life sciences, chemistry\u002Fmaterials, and physical world models. This topic covers the key milestones: UniPert unifying genetic-chemical perturbation spaces, GPT-Rosalind's end-to-end life-sciences reasoning, DAMO's AI agent discovering 4 superconducting materials in 28 hours, and Anthropic's Claude Science workbench going mainstream.","988a4300-5fab-41c4-b5d8-63711a2dc757","2026-09-10T01:34:15.296649Z","2026-09-10T01:34:15.296663Z",{"slug":49,"tag_slug":49,"title_zh":50,"title_en":51,"intro_zh":52,"intro_en":53,"id":54,"is_active":35,"created_at":55,"modified_at":56},"h3-series","MiniMax H3 系列：从开源权重到 35 倍吞吐","MiniMax H3 Series: from open weights to 35x throughput","MiniMax H3 自 2026 年 8 月开源以来节奏密集：官方把生成、参考与编辑收回一个模型；ComfyUI 当天压进 RTX 3060；摩尔线程 3 小时完成国产 GPU 适配；fal 后训练版把吞吐拉到 35 倍；FastH3 蒸馏再砍推理成本。本专题持续追踪 H3 的发布—开源—蒸馏—部署全链路。","Since MiniMax open-sourced H3 in August 2026 the pace has been relentless: one unified omni-modal model, same-day ComfyUI support down to an RTX 3060, a 3-hour Day-0 port to Moore Threads GPUs, fal's post-trained H3 Max at 35x throughput, and FastH3 distillation cutting inference cost further. This topic tracks the full H3 chain — release, open weights, distillation, deployment.","83ef0daa-3c31-4cb3-86ed-e5ee58654d5f","2026-09-08T07:33:19.942193Z","2026-09-08T07:33:19.942209Z",{"items":58},[59,64,69,74,79,84],{"id":60,"title":61,"news_slug":62,"published_at":63},"58ed753e-ad6d-4aac-95f4-36bf217e169c","把 10 万条人类视频变成机器人教材:RoboTok 检索 mAP 提升约 50 倍,hard 任务 79.3% 对 19.5%","robotok-retrieval-benchmark-reread","2026-09-06T21:11:25+00:00",{"id":65,"title":66,"news_slug":67,"published_at":68},"4c4a2a9e-f69b-4985-bd42-97ab2ef4e2ac","Spark-X2.5-4B 开源:4B 跑 1M 上下文,22 项基准打 9B 级 Qwen3.5","spark-x2-5-4b-apache-open-source","2026-09-16T01:30:00+00:00",{"id":70,"title":71,"news_slug":72,"published_at":73},"30fca629-bace-4832-9789-b44aa8c8989d","学生团队从零训出开源 7B 模型 ZGCM-1:数学推理硬刚 235B 前沿","zgcm-1-open-7b-foundation-model","2026-09-15T19:10:00+00:00",{"id":75,"title":76,"news_slug":77,"published_at":78},"54b86d93-0fd0-4107-9353-9b79a1446f69","NVIDIA 开源 IMO 金牌完整配方:30\u002F42 分、561B 双专家、算力账本全公开","nvidia-nemotron-imo-gold-open-recipe","2026-09-11T17:13:27+00:00",{"id":80,"title":81,"news_slug":82,"published_at":83},"2638aeac-dc4d-4b73-b7fe-2b042015adee","OreoLook 开源:三层缓存把 AI 搜索搬进 8 核 CPU,重复问题 0.1 毫秒出答案","oreolook-three-layer-cpu-cache","2026-09-10T23:08:36+00:00",{"id":85,"title":86,"news_slug":87,"published_at":88},"d41175a7-ad10-4e00-9017-a148fa0a77b3","BenchMIRT 把 LLM 基准拆到单题:Ai2 想让模型排名不再「一张考卷定生死」","ai2-benchmirt-llm-benchmark-audit","2026-09-10T11:05:05+00:00"]