Summary (Overview)

  • TokenRouter is a novel serving system designed for token-level LLM routing, where different tokens within a single response can be generated by different models.
  • The system introduces a request-centric programming, model-centric execution paradigm, decoupling routing algorithm development from serving-system optimization.
  • Key innovations include a route-send-receive programming interface, asynchronous tri-loop execution with a handoff-resume mechanism, and a delayed-batching scheduler with throughput-optimal hyperparameters derived from a Discrete-Time Markov Chain (DTMC) model.
  • Across 15 algorithm–workload combinations, TokenRouter achieves 2.01–64.15× higher decoding throughput than existing systems, and 2.73–21.97× over official implementations under original algorithmic settings.
  • TokenRouter advances the throughput-accuracy Pareto frontier, making token-level routing competitive with query-level routing on state-of-the-art frameworks like SGLang.

Introduction and Theoretical Foundation

Background and Motivation

Large language models (LLMs) vary in size, latency, and capability. Model routing exploits this diversity to improve the cost-quality Pareto frontier of LLM inference. While coarse-grained routing (session- or query-level) is widely deployed in production systems (e.g., ChatGPT, Cursor, RouteLLM), it binds an entire request to a single model.

Token-level routing offers two key advantages over query-level routing:

  1. Efficiency: It exploits variation in generation difficulty within a single query. For example, R2R can match a 32B model's quality while routing only 5% of tokens to the 32B model and decoding the rest with a 1.5B model (versus ~40% of queries routed to the 32B model at query level).
  2. Quality: Models with complementary expertise can collaborate within a single response, potentially surpassing any single LLM's quality, while providing smooth control over the cost-quality trade-off.

Challenges with Existing Systems

Existing serving systems (SGLang, vLLM) are designed for single-LLM serving, where all active requests are synchronous at each decoding step. Token-level routing breaks this assumption, creating three main challenges:

  1. Step desynchronization: Different LLMs have widely different per-step latencies. Synchronized batches force every step to wait for the slowest model, leaving faster models idle.
  2. Batch admission delay: Frequent model switches mean a routed request often arrives while its target model is still processing a previous batch, creating bubbles and fragmenting batches.
  3. Implementation complexity: No programming interface exists for per-step routing decisions, requiring extensive modifications to large codebases.

Theoretical Foundation

TokenRouter models the routing process as a Discrete-Time Markov Chain (DTMC). Given a token-level routing algorithm, concurrency NN, routing probability PP, and per-step decoding latency LiL_i of each LLM, the model derives throughput as a function of the batching threshold BB and searches for the optimal threshold B∗B^* that maximizes throughput.


Methodology

Programming Interface: Route-Send-Receive

TokenRouter abstracts diverse token-level routing algorithms with three components:

  • route(result): Called after each decoding step. Determines a destination index for each request based on forward-pass results (hidden states, logits, sampled tokens). Destination 0 = continue locally; nonzero = delegate to peer model.
  • send(req): Called when route delegates a request. Returns a PeerReq message carrying the request ID, token suffix unseen by the peer, a status field, and algorithm-specific fields. A position anchor req.loc tracks the decoding position.
  • receive(peer_req): Called when a PeerReq arrives. Converts it to the local request format. Default behavior commits all transferred tokens; developers override for algorithm-specific handoff semantics.

The lifecycle is expressed as:

receive→decode→route→send→peer receive→peer decode→peer route→⋯\text{receive} \rightarrow \text{decode} \rightarrow \text{route} \rightarrow \text{send} \rightarrow \text{peer receive} \rightarrow \text{peer decode} \rightarrow \text{peer route} \rightarrow \cdots

System Architecture

TokenRouter organizes the runtime as multiple subservers behind a single external interface:

  • Each subserver hosts one candidate LLM with its own scheduler, LLM runner, private KV-cache pool, and user-defined routing functions.
  • Subservers progress independently and communicate through peer requests.
  • The external interface exposes a single endpoint, making TokenRouter a drop-in replacement for standard single-LLM servers.

Asynchronous Tri-Loop Execution

TokenRouter extends the standard two-loop structure (client-server loop and decoding loop) with a third inter-model loop:

  • Client-server loop (purple): Request admission and response streaming.
  • Decoding loop (blue): Request scheduling and model execution.
  • Inter-model loop (black): Sending and receiving peer requests between subservers.

This allows a routed request to leave the local batch and resume when the peer request returns, while remaining local requests continue executing at the local decoding loop's pace.

Handoff-Resume Mechanism

A pending state is introduced between the standard running and finished states:

  • A pending request is skipped by the scheduler but its serving state is preserved.
  • At handoff, the routed request is marked pending.
  • Upon resumption, pending toggles to running or finished based on peer output.
  • New tokens are appended without redoing prefix matching or KV-cache allocation.

Delayed-Batching Scheduler

The scheduler buffers received requests and launches a batch only when the buffer size reaches threshold BB. This trades a small waiting time for larger same-model batches:

  • Too small BB → large batch admission delay.
  • Too large BB → requests trapped in queue, starving models of active work.

The throughput-optimal threshold B∗B^* is derived analytically from the DTMC model.


Empirical Validation / Results

Experimental Setup

  • Models: Qwen3-0.6B, Qwen3-8B, Qwen3-32B on 8×A100-80G GPUs.
  • Algorithms: CITER, R2R, R-Stitch, Co-LLM (two-model), and ME (three-model ensemble).
  • Baselines: Official released code and a standard SGLang-based baseline (Std. Serving).
  • Workloads: Low-effort reasoning (short I/O), high-effort reasoning (short input, long output), and agentic tasks (multi-turn, long input).

Serving Efficiency Results

AlgorithmImplementationThroughput (token/s)TTFT (s)Latency (s)
R2RLLM-only145.300.13411.98
R2ROfficial Code89.620.11751.15
R2RTokenRouter244.560.11270.19
CITERLLM-only123.720.0832.68
CITEROfficial Code17.160.0367.64
CITERTokenRouter149.310.0360.48
Co-LLMLLM-only134.790.06913.39
Co-LLMOfficial Code3.461.14247.61
Co-LLMTokenRouter76.020.06711.46
R-StitchLLM-only150.410.15360.76
R-StitchTokenRouter140.580.067151.30

Key findings:

  • Across all 15 algorithm–workload combinations, TokenRouter improves throughput by 2.01–64.15× over the stronger baseline.
  • TokenRouter reduces end-to-end latency of Std. Serving by 2.03–63.64×.
  • TokenRouter scales gracefully with output length: throughput holds steady while Std. Serving loses 58.1–85.2% of throughput on long outputs.

Model Pair Generalization

SLMLLMR2R (token/s)TokenRouter (token/s)Gain
0.6B8B105.1336.983.21×
0.6B32B77.08197.402.56×
1.7B8B103.37285.412.76×
4B8B106.67211.811.99×

Ablation Study

At concurrency 8:

  • Engineering optimizations (extended CUDA graphs): 132.78 → 230.79 token/s (1.71× over official R2R).
  • Asynchronous execution: → 296.86 token/s.
  • Delayed batching: → 372.48 token/s (2.76× overall).

Throughput–Speed Trade-off

  • As concurrency increases from 1 to 16, TokenRouter's throughput increases by 8.61× while retaining 51.7% of single-user speed.
  • At concurrency 16, TokenRouter achieves 18.58× higher throughput than official R2R at concurrency 1, while still providing 1.13× higher per-user speed (i.e., better throughput even under a stricter SLO).

Pareto Frontier

TokenRouter shifts token-level routing to a new throughput-accuracy Pareto frontier, making fine-grained routing more competitive than query-level routing (RouteLLM) on SGLang.


Theoretical and Practical Implications

Theoretical Contributions

  1. DTMC-based throughput model: Provides a principled framework for deriving optimal batching thresholds in token-routing systems, balancing admission delay against batch fragmentation.
  2. Formalization of routing lifecycle: The receive → decode → route → send abstraction generalizes across diverse token-level routing algorithms (confidence-based, entropy-based, ensemble-based, deferral-based).
  3. Separation of concerns: The request-centric programming / model-centric execution principle decouples algorithmic design from systems optimization.

Practical Implications

  1. Drop-in replacement: TokenRouter's single external interface means existing clients can use it without modification, regardless of how many LLMs participate in routing.
  2. Developer productivity: The three-function interface (route, send, receive) dramatically reduces implementation complexity compared to modifying single-LLM serving systems.
  3. Production viability: Token-level routing becomes practically deployable, unlocking cost savings (e.g., matching 32B quality with only 5% of tokens routed to the large model) that were previously theoretical.
  4. Scalability: The identical subserver layout makes it easy to scale to more than two LLMs (demonstrated with ME, a three-model ensemble).

Conclusion

TokenRouter is an efficient and developer-friendly serving system for token-level LLM routing. Its key contributions are:

  1. Request-centric programming interface (route-send-receive) that lets developers express diverse routing algorithms simply.
  2. Asynchronous tri-loop execution with handoff-resume that eliminates step desynchronization across cooperating LLMs.
  3. Delayed-batching scheduler with analytically derived optimal thresholds that reduces batch admission delay.

Across diverse routing algorithms, workloads, and model pairs, TokenRouter consistently achieves 2.01–64.15× higher decoding throughput than existing implementations, providing a strong foundation for future token-routing research.

Limitations and Future Work

The mathematical model assumes the number of tokens between two consecutive sends follows a geometric distribution, which holds for most token-level routing algorithms. Handling corner cases that violate this assumption (e.g., algorithms with periodic or state-dependent switching patterns) is left for future work.

Related papers