Within a single response, let a small cheap model generate most tokens and escalate only the hard ones to a large model — that is the token-level routing recipe for cutting inference cost. The algorithm side has been busy for a while, and the paper's abstract states it plainly: coarse-grained session/query-level routing is already widely adopted in production, while recent algorithmic work shows token-level routing delivers substantial efficiency and quality gains. The missing piece has been the system side — existing inference engines are built on the assumption that one request stays bound to one model from start to finish.
Where existing systems stall
The paper breaks the problem into three layers. First, step desynchronization: two models decode at naturally different speeds, and the faster one stalls waiting for the slower peer. Second, batch admission delays: tokens travel between models irregularly, so conventional continuous batching cannot form effective batches. Third, implementation complexity: routing-algorithm authors must hand-roll batch scheduling, model handoff, and cache state, which keeps these algorithms stuck in papers.
TokenRouter's answer
The team from NICS-EFC at Tsinghua's Department of Electronic Engineering calls it TokenRouter: accepted at NeurIPS 2026, paper on arXiv in October, code open-sourced alongside. The design principle in one line: request-centric programming, model-centric execution — developers describe per-request routing logic through three interfaces, route(), send(), and receive(), while the runtime launches an independent subserver for each model, each running its own decoding loop, dispatched asynchronously.
Two engineering details stand out. First, only token suffixes and routing state travel between models; KV caches stay local and are never migrated. A routed-out request stays pending and keeps its serving state and KV slot; returning tokens are appended directly, with no repeated prefix matching or KV allocation. Second, each subserver runs a delayed-batching scheduler that gathers irregular token arrivals into batches, and the optimal threshold is not hand-tuned — it is derived from a mathematical throughput model of the system.
Developer ergonomics got attention too: the system extends SGLang's server args, is configured through a single YAML file, supports placing one model per node across two nodes, exposes an OpenAI-compatible endpoint, and can also be embedded directly in Python via TokenRoutingEngine.
Numbers and ecosystem position
The official numbers from the paper: across diverse routing algorithms, workloads, and model pairs, decoding throughput is 2.01x to 64.15x higher than existing systems. Five routing algorithms are covered — CITER, R2R, R-Stitch, Co-LLM, and ME — plus GlimpRouter, query-level routing, and a random baseline. Evaluated model pairs include Qwen2 1.5B/72B, LLaMA2 7B/70B, and L1-1.5B-short/QwQ-32B, across CommonsenseQA, GSM8K, and AIME. R2R itself is the same group's earlier algorithm work — this release is effectively the engine for their own algorithm family.
My take
The real value is not the 64x ceiling but turning "switch routing policy" from "modify the system" into "swap a YAML file". Token-level routing algorithms appear several times a year; without a serving engine, they remain tables inside papers. Now the algorithm-system loop is closed. Two buckets of cold water: 64.15x is the upper end of a range whose floor is 2.01x, and throughput gains are not quality gains — response quality still depends on the routing algorithm's own judgment. The repo sits at 20 stars and 2 commits, early cold-start days — worth watching, not yet production material.
References: arXiv 2610.12242 (https://arxiv.org/abs/2610.12242); GitHub thu-nics/TokenRouter (https://github.com/thu-nics/TokenRouter)