Summary (Overview)
- TokenRouter is a novel serving system designed for token-level LLM routing, where different tokens within a single response can be generated by different models.
- The system introduces a request-centric programming, model-centric execution paradigm, decoupling routing algorithm development from serving-system optimization.
- Key innovations include a route-send-receive programming interface, asynchronous tri-loop execution with a handoff-resume mechanism, and a delayed-batching scheduler with throughput-optimal hyperparameters derived from a Discrete-Time Markov Chain (DTMC) model.
- Across 15 algorithm–workload combinations, TokenRouter achieves 2.01–64.15× higher decoding throughput than existing systems, and 2.73–21.97× over official implementations under original algorithmic settings.
- TokenRouter advances the throughput-accuracy Pareto frontier, making token-level routing competitive with query-level routing on state-of-the-art frameworks like SGLang.
Introduction and Theoretical Foundation
Background and Motivation
Large language models (LLMs) vary in size, latency, and capability. Model routing exploits this diversity to improve the cost-quality Pareto frontier of LLM inference. While coarse-grained routing (session- or query-level) is widely deployed in production systems (e.g., ChatGPT, Cursor, RouteLLM), it binds an entire request to a single model.
Token-level routing offers two key advantages over query-level routing:
- Efficiency: It exploits variation in generation difficulty within a single query. For example, R2R can match a 32B model's quality while routing only 5% of tokens to the 32B model and decoding the rest with a 1.5B model (versus ~40% of queries routed to the 32B model at query level).
- Quality: Models with complementary expertise can collaborate within a single response, potentially surpassing any single LLM's quality, while providing smooth control over the cost-quality trade-off.
Challenges with Existing Systems
Existing serving systems (SGLang, vLLM) are designed for single-LLM serving, where all active requests are synchronous at each decoding step. Token-level routing breaks this assumption, creating three main challenges:
- Step desynchronization: Different LLMs have widely different per-step latencies. Synchronized batches force every step to wait for the slowest model, leaving faster models idle.
- Batch admission delay: Frequent model switches mean a routed request often arrives while its target model is still processing a previous batch, creating bubbles and fragmenting batches.
- Implementation complexity: No programming interface exists for per-step routing decisions, requiring extensive modifications to large codebases.
Theoretical Foundation
TokenRouter models the routing process as a Discrete-Time Markov Chain (DTMC). Given a token-level routing algorithm, concurrency , routing probability , and per-step decoding latency of each LLM, the model derives throughput as a function of the batching threshold and searches for the optimal threshold that maximizes throughput.
Methodology
Programming Interface: Route-Send-Receive
TokenRouter abstracts diverse token-level routing algorithms with three components:
route(result): Called after each decoding step. Determines a destination index for each request based on forward-pass results (hidden states, logits, sampled tokens). Destination 0 = continue locally; nonzero = delegate to peer model.send(req): Called whenroutedelegates a request. Returns aPeerReqmessage carrying the request ID, token suffix unseen by the peer, a status field, and algorithm-specific fields. A position anchorreq.loctracks the decoding position.receive(peer_req): Called when aPeerReqarrives. Converts it to the local request format. Default behavior commits all transferred tokens; developers override for algorithm-specific handoff semantics.
The lifecycle is expressed as:
System Architecture
TokenRouter organizes the runtime as multiple subservers behind a single external interface:
- Each subserver hosts one candidate LLM with its own scheduler, LLM runner, private KV-cache pool, and user-defined routing functions.
- Subservers progress independently and communicate through peer requests.
- The external interface exposes a single endpoint, making TokenRouter a drop-in replacement for standard single-LLM servers.
Asynchronous Tri-Loop Execution
TokenRouter extends the standard two-loop structure (client-server loop and decoding loop) with a third inter-model loop:
- Client-server loop (purple): Request admission and response streaming.
- Decoding loop (blue): Request scheduling and model execution.
- Inter-model loop (black): Sending and receiving peer requests between subservers.
This allows a routed request to leave the local batch and resume when the peer request returns, while remaining local requests continue executing at the local decoding loop's pace.
Handoff-Resume Mechanism
A pending state is introduced between the standard running and finished states:
- A pending request is skipped by the scheduler but its serving state is preserved.
- At handoff, the routed request is marked pending.
- Upon resumption, pending toggles to running or finished based on peer output.
- New tokens are appended without redoing prefix matching or KV-cache allocation.
Delayed-Batching Scheduler
The scheduler buffers received requests and launches a batch only when the buffer size reaches threshold . This trades a small waiting time for larger same-model batches:
- Too small → large batch admission delay.
- Too large → requests trapped in queue, starving models of active work.
The throughput-optimal threshold is derived analytically from the DTMC model.
Empirical Validation / Results
Experimental Setup
- Models: Qwen3-0.6B, Qwen3-8B, Qwen3-32B on 8×A100-80G GPUs.
- Algorithms: CITER, R2R, R-Stitch, Co-LLM (two-model), and ME (three-model ensemble).
- Baselines: Official released code and a standard SGLang-based baseline (Std. Serving).
- Workloads: Low-effort reasoning (short I/O), high-effort reasoning (short input, long output), and agentic tasks (multi-turn, long input).
Serving Efficiency Results
| Algorithm | Implementation | Throughput (token/s) | TTFT (s) | Latency (s) |
|---|---|---|---|---|
| R2R | LLM-only | 145.30 | 0.13 | 411.98 |
| R2R | Official Code | 89.62 | 0.11 | 751.15 |
| R2R | TokenRouter | 244.56 | 0.11 | 270.19 |
| CITER | LLM-only | 123.72 | 0.083 | 2.68 |
| CITER | Official Code | 17.16 | 0.036 | 7.64 |
| CITER | TokenRouter | 149.31 | 0.036 | 0.48 |
| Co-LLM | LLM-only | 134.79 | 0.069 | 13.39 |
| Co-LLM | Official Code | 3.46 | 1.14 | 247.61 |
| Co-LLM | TokenRouter | 76.02 | 0.067 | 11.46 |
| R-Stitch | LLM-only | 150.41 | 0.15 | 360.76 |
| R-Stitch | TokenRouter | 140.58 | 0.067 | 151.30 |
Key findings:
- Across all 15 algorithm–workload combinations, TokenRouter improves throughput by 2.01–64.15× over the stronger baseline.
- TokenRouter reduces end-to-end latency of Std. Serving by 2.03–63.64×.
- TokenRouter scales gracefully with output length: throughput holds steady while Std. Serving loses 58.1–85.2% of throughput on long outputs.
Model Pair Generalization
| SLM | LLM | R2R (token/s) | TokenRouter (token/s) | Gain |
|---|---|---|---|---|
| 0.6B | 8B | 105.1 | 336.98 | 3.21× |
| 0.6B | 32B | 77.08 | 197.40 | 2.56× |
| 1.7B | 8B | 103.37 | 285.41 | 2.76× |
| 4B | 8B | 106.67 | 211.81 | 1.99× |
Ablation Study
At concurrency 8:
- Engineering optimizations (extended CUDA graphs): 132.78 → 230.79 token/s (1.71× over official R2R).
- Asynchronous execution: → 296.86 token/s.
- Delayed batching: → 372.48 token/s (2.76× overall).
Throughput–Speed Trade-off
- As concurrency increases from 1 to 16, TokenRouter's throughput increases by 8.61× while retaining 51.7% of single-user speed.
- At concurrency 16, TokenRouter achieves 18.58× higher throughput than official R2R at concurrency 1, while still providing 1.13× higher per-user speed (i.e., better throughput even under a stricter SLO).
Pareto Frontier
TokenRouter shifts token-level routing to a new throughput-accuracy Pareto frontier, making fine-grained routing more competitive than query-level routing (RouteLLM) on SGLang.
Theoretical and Practical Implications
Theoretical Contributions
- DTMC-based throughput model: Provides a principled framework for deriving optimal batching thresholds in token-routing systems, balancing admission delay against batch fragmentation.
- Formalization of routing lifecycle: The receive → decode → route → send abstraction generalizes across diverse token-level routing algorithms (confidence-based, entropy-based, ensemble-based, deferral-based).
- Separation of concerns: The request-centric programming / model-centric execution principle decouples algorithmic design from systems optimization.
Practical Implications
- Drop-in replacement: TokenRouter's single external interface means existing clients can use it without modification, regardless of how many LLMs participate in routing.
- Developer productivity: The three-function interface (route, send, receive) dramatically reduces implementation complexity compared to modifying single-LLM serving systems.
- Production viability: Token-level routing becomes practically deployable, unlocking cost savings (e.g., matching 32B quality with only 5% of tokens routed to the large model) that were previously theoretical.
- Scalability: The identical subserver layout makes it easy to scale to more than two LLMs (demonstrated with ME, a three-model ensemble).
Conclusion
TokenRouter is an efficient and developer-friendly serving system for token-level LLM routing. Its key contributions are:
- Request-centric programming interface (route-send-receive) that lets developers express diverse routing algorithms simply.
- Asynchronous tri-loop execution with handoff-resume that eliminates step desynchronization across cooperating LLMs.
- Delayed-batching scheduler with analytically derived optimal thresholds that reduces batch admission delay.
Across diverse routing algorithms, workloads, and model pairs, TokenRouter consistently achieves 2.01–64.15× higher decoding throughput than existing implementations, providing a strong foundation for future token-routing research.
Limitations and Future Work
The mathematical model assumes the number of tokens between two consecutive sends follows a geometric distribution, which holds for most token-level routing algorithms. Handling corner cases that violate this assumption (e.g., algorithms with periodic or state-dependent switching patterns) is left for future work.
Related papers
- STEPQuant: When and Where Errors Matter in Delta-Rule Recurrent State Quantization
STEPQuant achieves near-FP32 accuracy at 6-bit recurrent state quantization via lifetime-aware bit allocation and key-row-aware dual-axis fitting, cutting serving memory by 68.7%.
- How Far Are We from Removing the Visual Encoder? Scaling Laws for Encoder-Free Multimodal Pretraining
Encoder-free multimodal LLMs match encoder-based performance at ~10^22 FLOPs, shifting compute-optimal allocation toward larger decoders and enabling viable encoder-free pretraining.
- How Much of a Harness Does a Strong Agent Need for Autonomous ML Engineering?
Under matched conditions, a minimal-harness coding agent matches or outperforms state-of-the-art MLE harnesses, with performance driven by the LLM backbone and execution environment, not scaffolding.