[{"data":1,"prerenderedAt":-1},["ShallowReactive",2],{"news-slug-tokenrouter-token-level-llm-routing":3,"topics-all":38,"news-related-95e663a4-2cdf-454c-9787-b154fbd41909":57},{"id":4,"title":5,"summary":6,"content":7,"original_url":8,"source_id":9,"tags":10,"translations":24,"news_slug":31,"published_at":32,"created_at":33,"modified_at":34,"is_published":35,"publish_type":36,"image_url":14,"view_count":37},"95e663a4-2cdf-454c-9787-b154fbd41909","TokenRouter:token 级路由提速 64 倍","token 级路由有了专用推理引擎。清华团队的 TokenRouter 给每个模型独立 subserver,只传 token 后缀与路由状态,KV 缓存不动,再用 delayed-batching 攒批。五种路由算法、多组模型上,解码吞吐较现有系统最高提升 64.15 倍,NeurIPS 2026 接收,代码开源。","一个回答里让小模型和大模型接力——绝大多数 token 由便宜的小模型生成,只在真正难的地方递给大模型——这是 token 级路由给出的推理降本路线。算法侧这条路已经热闹了一阵,论文摘要里也直言:粗粒度的 session\u002Fquery 级路由在生产系统里已被广泛采用,而近期算法工作显示 token 级路由能带来可观的效率与质量收益。短板一直卡在系统侧——现有推理引擎是按「一个请求从头到尾绑死一个模型」的假设设计的。\n\n## 卡在哪\n\n论文把问题拆成三层。其一,步失步(step desynchronization):两个模型解码速度天然不同,快的一方要停下来等慢的一方。其二,批准入延迟:token 在模型间不规则往返,常规连续批处理凑不出有效批次。其三,实现复杂度:路由算法作者要亲自处理批调度、模型交接和缓存状态,算法很难走出论文。\n\n## TokenRouter 的答案\n\n清华电子系 NICS-EFC 团队的方案叫 TokenRouter,已被 NeurIPS 2026 接收,论文 10 月上 arXiv,代码同步开源。设计原则一句话:**请求中心编程,模型中心执行**——开发者用 route()、send()、receive() 三个接口只描述单个请求的路由逻辑;运行时为每个模型启动一个独立的 subserver,各自维持解码循环、异步调度。\n\n两个工程细节值得单说。一是模型之间只传 token 后缀和路由状态,KV 缓存全部留在本地、不搬运;被路由出去的请求保持 pending 状态并保住自己的 KV 槽位,token 回流后直接续写,不重做前缀匹配和缓存分配。二是每个 subserver 配一个 delayed-batching 调度器,把不规则的 token 到达攒成批次,而最优阈值不是拍脑袋定的,由团队给系统建的数学吞吐模型推导出来。\n\n工程亲和度也照顾到了:系统基于 SGLang 的 server args 扩展,单 YAML 文件完成配置,支持双节点各放一个模型,对外提供 OpenAI 兼容接口,也可以用 TokenRoutingEngine 直接在 Python 里内嵌使用。\n\n## 数字与生态位\n\n论文报告的官方数字:在多种路由算法、负载和模型组合上,解码吞吐比现有系统高 2.01 到 64.15 倍。覆盖的路由算法有五种——CITER、R2R、R-Stitch、Co-LLM、ME,另有 GlimpRouter、query 级路由和随机基线;实测模型对包括 Qwen2 1.5B\u002F72B、LLaMA2 7B\u002F70B、L1-1.5B-short\u002FQwQ-32B 等,横跨 CommonsenseQA、GSM8K、AIME 三类基准。其中 R2R 本身就是同组的前序算法工作,这次相当于给自己的算法家族补上了发动机。\n\n## 我的看法\n\n这件事的价值不在 64 倍这个区间上限,而在把「换路由策略」从改系统变成换一个 YAML 文件。token 级路由算法一年能出好几个,但没有 serving 引擎承接,它们就只是论文里的表格;现在算法与系统的闭环补上了。冷水也要泼两条:64.15x 是极端配置下的上限,区间下限是 2.01x;且吞吐提升不等于质量提升,回答好坏仍取决于路由算法自身的判断力。repo 目前 20 star、2 个 commit,还在冷启动期——值得盯,但别急着上生产。\n\n参考:arXiv 2610.12242(https:\u002F\u002Farxiv.org\u002Fabs\u002F2610.12242);GitHub thu-nics\u002FTokenRouter(https:\u002F\u002Fgithub.com\u002Fthu-nics\u002FTokenRouter)","https:\u002F\u002Farxiv.org\u002Fabs\u002F2610.12242","7437aeb9-930c-4866-a2e9-48003c1a792b",[11,15,18,21],{"id":12,"name":13,"slug":13,"description":14,"color":14},"fca9258a-9430-455a-b95d-b9fae5e373a8","ai-inference",null,{"id":16,"name":17,"slug":17,"description":14,"color":14},"0ef8513a-0a26-42f0-b6f9-5b6dadded45c","efficiency",{"id":19,"name":20,"slug":20,"description":14,"color":14},"01598627-1ea6-4b27-a5d8-874971571a71","llm",{"id":22,"name":23,"slug":23,"description":14,"color":14},"b9bd9039-fcdb-41a8-b85b-fc1587def2b9","open-source",[25],{"id":26,"lang":27,"title":28,"summary":29,"content":30},"b4f423b3-aee3-4b39-ae49-58ee578ef11f","en","TokenRouter: token-level LLM routing gets a serving engine","Tsinghua's TokenRouter: per-model subservers + delayed batching serve token-level LLM routing at 2.01-64.15x throughput. NeurIPS 2026, open source.","Within a single response, let a small cheap model generate most tokens and escalate only the hard ones to a large model — that is the token-level routing recipe for cutting inference cost. The algorithm side has been busy for a while, and the paper's abstract states it plainly: coarse-grained session\u002Fquery-level routing is already widely adopted in production, while recent algorithmic work shows token-level routing delivers substantial efficiency and quality gains. The missing piece has been the system side — existing inference engines are built on the assumption that one request stays bound to one model from start to finish.\n\n## Where existing systems stall\n\nThe paper breaks the problem into three layers. First, step desynchronization: two models decode at naturally different speeds, and the faster one stalls waiting for the slower peer. Second, batch admission delays: tokens travel between models irregularly, so conventional continuous batching cannot form effective batches. Third, implementation complexity: routing-algorithm authors must hand-roll batch scheduling, model handoff, and cache state, which keeps these algorithms stuck in papers.\n\n## TokenRouter's answer\n\nThe team from NICS-EFC at Tsinghua's Department of Electronic Engineering calls it TokenRouter: accepted at NeurIPS 2026, paper on arXiv in October, code open-sourced alongside. The design principle in one line: **request-centric programming, model-centric execution** — developers describe per-request routing logic through three interfaces, route(), send(), and receive(), while the runtime launches an independent subserver for each model, each running its own decoding loop, dispatched asynchronously.\n\nTwo engineering details stand out. First, only token suffixes and routing state travel between models; KV caches stay local and are never migrated. A routed-out request stays pending and keeps its serving state and KV slot; returning tokens are appended directly, with no repeated prefix matching or KV allocation. Second, each subserver runs a delayed-batching scheduler that gathers irregular token arrivals into batches, and the optimal threshold is not hand-tuned — it is derived from a mathematical throughput model of the system.\n\nDeveloper ergonomics got attention too: the system extends SGLang's server args, is configured through a single YAML file, supports placing one model per node across two nodes, exposes an OpenAI-compatible endpoint, and can also be embedded directly in Python via TokenRoutingEngine.\n\n## Numbers and ecosystem position\n\nThe official numbers from the paper: across diverse routing algorithms, workloads, and model pairs, decoding throughput is 2.01x to 64.15x higher than existing systems. Five routing algorithms are covered — CITER, R2R, R-Stitch, Co-LLM, and ME — plus GlimpRouter, query-level routing, and a random baseline. Evaluated model pairs include Qwen2 1.5B\u002F72B, LLaMA2 7B\u002F70B, and L1-1.5B-short\u002FQwQ-32B, across CommonsenseQA, GSM8K, and AIME. R2R itself is the same group's earlier algorithm work — this release is effectively the engine for their own algorithm family.\n\n## My take\n\nThe real value is not the 64x ceiling but turning \"switch routing policy\" from \"modify the system\" into \"swap a YAML file\". Token-level routing algorithms appear several times a year; without a serving engine, they remain tables inside papers. Now the algorithm-system loop is closed. Two buckets of cold water: 64.15x is the upper end of a range whose floor is 2.01x, and throughput gains are not quality gains — response quality still depends on the routing algorithm's own judgment. The repo sits at 20 stars and 2 commits, early cold-start days — worth watching, not yet production material.\n\nReferences: arXiv 2610.12242 (https:\u002F\u002Farxiv.org\u002Fabs\u002F2610.12242); GitHub thu-nics\u002FTokenRouter (https:\u002F\u002Fgithub.com\u002Fthu-nics\u002FTokenRouter)","tokenrouter-token-level-llm-routing","2026-10-09T21:11:28Z","2026-10-09T21:12:54.435962Z","2026-10-09T21:12:54.435978Z",true,"agent",3,[39,48],{"slug":40,"tag_slug":40,"title_zh":41,"title_en":42,"intro_zh":43,"intro_en":44,"id":45,"is_active":35,"created_at":46,"modified_at":47},"ai-for-science","AI for Science 2026：从 UniPert 到 GPT-Rosalind 的硬核进化","AI for Science 2026: from UniPert to GPT-Rosalind","生命科学、化学材料、物理世界模型——AI 正在从\"语言工具\"变成\"实验伙伴\"。本专题收录 AI 在三大科学方向的关键节点：UniPert 统一基因与化学扰动空间、GPT-Rosalind 端到端生命科学推理、达摩院 AI 智能体 28 小时找到 4 种超导新材料、Anthropic Claude Science 把工作台做成标准品。","From language tool to lab partner — AI is reshaping life sciences, chemistry\u002Fmaterials, and physical world models. This topic covers the key milestones: UniPert unifying genetic-chemical perturbation spaces, GPT-Rosalind's end-to-end life-sciences reasoning, DAMO's AI agent discovering 4 superconducting materials in 28 hours, and Anthropic's Claude Science workbench going mainstream.","988a4300-5fab-41c4-b5d8-63711a2dc757","2026-09-10T01:34:15.296649Z","2026-09-10T01:34:15.296663Z",{"slug":49,"tag_slug":49,"title_zh":50,"title_en":51,"intro_zh":52,"intro_en":53,"id":54,"is_active":35,"created_at":55,"modified_at":56},"h3-series","MiniMax H3 系列：从开源权重到 35 倍吞吐","MiniMax H3 Series: from open weights to 35x throughput","MiniMax H3 自 2026 年 8 月开源以来节奏密集：官方把生成、参考与编辑收回一个模型；ComfyUI 当天压进 RTX 3060；摩尔线程 3 小时完成国产 GPU 适配；fal 后训练版把吞吐拉到 35 倍；FastH3 蒸馏再砍推理成本。本专题持续追踪 H3 的发布—开源—蒸馏—部署全链路。","Since MiniMax open-sourced H3 in August 2026 the pace has been relentless: one unified omni-modal model, same-day ComfyUI support down to an RTX 3060, a 3-hour Day-0 port to Moore Threads GPUs, fal's post-trained H3 Max at 35x throughput, and FastH3 distillation cutting inference cost further. This topic tracks the full H3 chain — release, open weights, distillation, deployment.","83ef0daa-3c31-4cb3-86ed-e5ee58654d5f","2026-09-08T07:33:19.942193Z","2026-09-08T07:33:19.942209Z",{"items":58},[59,64,69,74,79,84],{"id":60,"title":61,"news_slug":62,"published_at":63},"ccce6dfe-776f-4cd8-9605-9163daea4627","Jeff 开源决策模型:2B 追平 Jev","jeff-open-decision-models-0-8b","2026-09-29T15:10:00+00:00",{"id":65,"title":66,"news_slug":67,"published_at":68},"523d3ae4-9c60-4bcb-8e6c-ea98e58c7f71","一次前向一个决策:GLM-5.3-Flash 平替 Jev","glm-flash-single-token-decisions","2026-09-28T13:11:55+00:00",{"id":70,"title":71,"news_slug":72,"published_at":73},"c6dc2edc-1a4a-46b9-85e6-f1c8ef32faa6","DeepSeek DSpark 跑进 Apple Silicon：mlx-dspark 给出首个原生 MLX 移植,逐字节保持原模型输出","mlx-dspark-apple-silicon","2026-07-04T12:00:00+00:00",{"id":75,"title":76,"news_slug":77,"published_at":78},"62e17707-e36f-45f6-8749-0d0370382cbd","llm-d：混合 GPU 集群 3-5 倍加速，KV Cache 感知路由","llm-d-mixed-gpu-kv-cache-aware-routing","2026-06-23T22:00:00+00:00",{"id":80,"title":81,"news_slug":82,"published_at":83},"267a9244-2ed7-4034-86cb-be4cbd196a08","NVIDIA 开源 Nemotron 3 Super：Latent MoE 如何让 120B 模型「省着跑」","nvidia-nemotron-3-super-120b-latent-moe","2026-05-18T04:05:00+00:00",{"id":85,"title":86,"news_slug":87,"published_at":88},"d6ec8624-ad4e-41ce-9566-22d1f0926d49","循环解码器+并行编码器:RLT长度外推翻盘","recurrent-looped-transformer-length-generalization","2026-10-08T23:30:00+00:00"]