# TokenRouter: Efficient Serving System for Token-Level LLM Routing

> TokenRouter is a serving system enabling token-level LLM routing with 2.01–64.15× higher throughput than existing implementations via asynchronous tri-loop execution and delayed-batching.

- **Source:** [arXiv](https://arxiv.org/abs/2610.12242)
- **Published:** 2026-10-10
- **Permalink:** https://picx.dev/p/oLGFPj
- **Whiteboard:** https://picx.dev/p/oLGFPj/image

## Summary

## Summary (Overview)

- **TokenRouter** is a novel serving system designed for token-level LLM routing, where different tokens within a single response can be generated by different models.
- The system introduces a **request-centric programming, model-centric execution** paradigm, decoupling routing algorithm development from serving-system optimization.
- Key innovations include a **route-send-receive programming interface**, **asynchronous tri-loop execution** with a handoff-resume mechanism, and a **delayed-batching scheduler** with throughput-optimal hyperparameters derived from a Discrete-Time Markov Chain (DTMC) model.
- Across 15 algorithm–workload combinations, TokenRouter achieves **2.01–64.15× higher decoding throughput** than existing systems, and 2.73–21.97× over official implementations under original algorithmic settings.
- TokenRouter advances the throughput-accuracy Pareto frontier, making token-level routing competitive with query-level routing on state-of-the-art frameworks like SGLang.

---

## Introduction and Theoretical Foundation

### Background and Motivation

Large language models (LLMs) vary in size, latency, and capability. **Model routing** exploits this diversity to improve the cost-quality Pareto frontier of LLM inference. While coarse-grained routing (session- or query-level) is widely deployed in production systems (e.g., ChatGPT, Cursor, RouteLLM), it binds an entire request to a single model.

**Token-level routing** offers two key advantages over query-level routing:

1. **Efficiency**: It exploits variation in generation difficulty *within* a single query. For example, R2R can match a 32B model's quality while routing only 5% of tokens to the 32B model and decoding the rest with a 1.5B model (versus ~40% of queries routed to the 32B model at query level).
2. **Quality**: Models with complementary expertise can collaborate within a single response, potentially surpassing any single LLM's quality, while providing smooth control over the cost-quality trade-off.

### Challenges with Existing Systems

Existing serving systems (SGLang, vLLM) are designed for **single-LLM serving**, where all active requests are synchronous at each decoding step. Token-level routing breaks this assumption, creating three main challenges:

1. **Step desynchronization**: Different LLMs have widely different per-step latencies. Synchronized batches force every step to wait for the slowest model, leaving faster models idle.
2. **Batch admission delay**: Frequent model switches mean a routed request often arrives while its target model is still processing a previous batch, creating bubbles and fragmenting batches.
3. **Implementation complexity**: No programming interface exists for per-step routing decisions, requiring extensive modifications to large codebases.

### Theoretical Foundation

TokenRouter models the routing process as a **Discrete-Time Markov Chain (DTMC)**. Given a token-level routing algorithm, concurrency $N$, routing probability $P$, and per-step decoding latency $L_i$ of each LLM, the model derives throughput as a function of the batching threshold $B$ and searches for the optimal threshold $B^*$ that maximizes throughput.

---

## Methodology

### Programming Interface: Route-Send-Receive

TokenRouter abstracts diverse token-level routing algorithms with three components:

- **`route(result)`**: Called after each decoding step. Determines a destination index for each request based on forward-pass results (hidden states, logits, sampled tokens). Destination 0 = continue locally; nonzero = delegate to peer model.
- **`send(req)`**: Called when `route` delegates a request. Returns a `PeerReq` message carrying the request ID, token suffix unseen by the peer, a status field, and algorithm-specific fields. A position anchor `req.loc` tracks the decoding position.
- **`receive(peer_req)`**: Called when a `PeerReq` arrives. Converts it to the local request format. Default behavior commits all transferred tokens; developers override for algorithm-specific handoff semantics.

The lifecycle is expressed as:

$$\text{receive} \rightarrow \text{decode} \rightarrow \text{route} \rightarrow \text{send} \rightarrow \text{peer receive} \rightarrow \text{peer decode} \rightarrow \text{peer route} \rightarrow \cdots$$

### System Architecture

TokenRouter organizes the runtime as **multiple subservers** behind a single external interface:

- Each subserver hosts one candidate LLM with its own scheduler, LLM runner, private KV-cache pool, and user-defined routing functions.
- Subservers progress independently and communicate through peer requests.
- The external interface exposes a single endpoint, making TokenRouter a **drop-in replacement** for standard single-LLM servers.

### Asynchronous Tri-Loop Execution

TokenRouter extends the standard two-loop structure (client-server loop and decoding loop) with a third **inter-model loop**:

- **Client-server loop (purple)**: Request admission and response streaming.
- **Decoding loop (blue)**: Request scheduling and model execution.
- **Inter-model loop (black)**: Sending and receiving peer requests between subservers.

This allows a routed request to leave the local batch and resume when the peer request returns, while remaining local requests continue executing at the local decoding loop's pace.

### Handoff-Resume Mechanism

A **pending state** is introduced between the standard `running` and `finished` states:

- A pending request is skipped by the scheduler but its serving state is preserved.
- At handoff, the routed request is marked pending.
- Upon resumption, pending toggles to running or finished based on peer output.
- New tokens are appended without redoing prefix matching or KV-cache allocation.

### Delayed-Batching Scheduler

The scheduler buffers received requests and launches a batch only when the buffer size reaches threshold $B$. This trades a small waiting time for larger same-model batches:

- Too small $B$ → large batch admission delay.
- Too large $B$ → requests trapped in queue, starving models of active work.

The **throughput-optimal threshold** $B^*$ is derived analytically from the DTMC model.

---

## Empirical Validation / Results

### Experimental Setup

- **Models**: Qwen3-0.6B, Qwen3-8B, Qwen3-32B on 8×A100-80G GPUs.
- **Algorithms**: CITER, R2R, R-Stitch, Co-LLM (two-model), and ME (three-model ensemble).
- **Baselines**: Official released code and a standard SGLang-based baseline (Std. Serving).
- **Workloads**: Low-effort reasoning (short I/O), high-effort reasoning (short input, long output), and agentic tasks (multi-turn, long input).

### Serving Efficiency Results

| Algorithm | Implementation | Throughput (token/s) | TTFT (s) | Latency (s) |
|---|---|---|---|---|
| R2R | LLM-only | 145.30 | 0.13 | 411.98 |
| R2R | Official Code | 89.62 | 0.11 | 751.15 |
| **R2R** | **TokenRouter** | **244.56** | **0.11** | **270.19** |
| CITER | LLM-only | 123.72 | 0.083 | 2.68 |
| CITER | Official Code | 17.16 | 0.036 | 7.64 |
| **CITER** | **TokenRouter** | **149.31** | **0.036** | **0.48** |
| Co-LLM | LLM-only | 134.79 | 0.069 | 13.39 |
| Co-LLM | Official Code | 3.46 | 1.14 | 247.61 |
| **Co-LLM** | **TokenRouter** | **76.02** | **0.067** | **11.46** |
| R-Stitch | LLM-only | 150.41 | 0.15 | 360.76 |
| **R-Stitch** | **TokenRouter** | **140.58** | **0.067** | **151.30** |

**Key findings**:
- Across all 15 algorithm–workload combinations, TokenRouter improves throughput by **2.01–64.15×** over the stronger baseline.
- TokenRouter reduces end-to-end latency of Std. Serving by **2.03–63.64×**.
- TokenRouter scales gracefully with output length: throughput holds steady while Std. Serving loses 58.1–85.2% of throughput on long outputs.

### Model Pair Generalization

| SLM | LLM | R2R (token/s) | TokenRouter (token/s) | Gain |
|---|---|---|---|---|
| 0.6B | 8B | 105.1 | 336.98 | 3.21× |
| 0.6B | 32B | 77.08 | 197.40 | 2.56× |
| 1.7B | 8B | 103.37 | 285.41 | 2.76× |
| 4B | 8B | 106.67 | 211.81 | 1.99× |

### Ablation Study

At concurrency 8:
- Engineering optimizations (extended CUDA graphs): 132.78 → 230.79 token/s (**1.71×** over official R2R).
- Asynchronous execution: → 296.86 token/s.
- Delayed batching: → 372.48 token/s (**2.76×** overall).

### Throughput–Speed Trade-off

- As concurrency increases from 1 to 16, TokenRouter's throughput increases by **8.61×** while retaining 51.7% of single-user speed.
- At concurrency 16, TokenRouter achieves **18.58× higher throughput** than official R2R at concurrency 1, while still providing 1.13× higher per-user speed (i.e., better throughput even under a stricter SLO).

### Pareto Frontier

TokenRouter shifts token-level routing to a new throughput-accuracy Pareto frontier, making fine-grained routing more competitive than query-level routing (RouteLLM) on SGLang.

---

## Theoretical and Practical Implications

### Theoretical Contributions

1. **DTMC-based throughput model**: Provides a principled framework for deriving optimal batching thresholds in token-routing systems, balancing admission delay against batch fragmentation.
2. **Formalization of routing lifecycle**: The receive → decode → route → send abstraction generalizes across diverse token-level routing algorithms (confidence-based, entropy-based, ensemble-based, deferral-based).
3. **Separation of concerns**: The request-centric programming / model-centric execution principle decouples algorithmic design from systems optimization.

### Practical Implications

1. **Drop-in replacement**: TokenRouter's single external interface means existing clients can use it without modification, regardless of how many LLMs participate in routing.
2. **Developer productivity**: The three-function interface (route, send, receive) dramatically reduces implementation complexity compared to modifying single-LLM serving systems.
3. **Production viability**: Token-level routing becomes practically deployable, unlocking cost savings (e.g., matching 32B quality with only 5% of tokens routed to the large model) that were previously theoretical.
4. **Scalability**: The identical subserver layout makes it easy to scale to more than two LLMs (demonstrated with ME, a three-model ensemble).

---

## Conclusion

TokenRouter is an efficient and developer-friendly serving system for token-level LLM routing. Its key contributions are:

1. **Request-centric programming interface** (route-send-receive) that lets developers express diverse routing algorithms simply.
2. **Asynchronous tri-loop execution** with handoff-resume that eliminates step desynchronization across cooperating LLMs.
3. **Delayed-batching scheduler** with analytically derived optimal thresholds that reduces batch admission delay.

Across diverse routing algorithms, workloads, and model pairs, TokenRouter consistently achieves **2.01–64.15× higher decoding throughput** than existing implementations, providing a strong foundation for future token-routing research.

### Limitations and Future Work

The mathematical model assumes the number of tokens between two consecutive sends follows a **geometric distribution**, which holds for most token-level routing algorithms. Handling corner cases that violate this assumption (e.g., algorithms with periodic or state-dependent switching patterns) is left for future work.

---

_Markdown view of https://picx.dev/p/oLGFPj, served by PicX — AI-generated visual whiteboard summaries of research papers._
