Beyond Accuracy and Cost: Latency-Aware LLM Query Routing for Dynamic Workloads
Quick Answer
This paper presents a latency-aware query routing system for language models, improving accuracy-cost utility by up to 40% while maintaining standard latencies.
Quick Take
The proposed lightweight latency estimator simulates token processing to optimize query assignments based on latency, accuracy, and cost, addressing the limitations of current latency-agnostic routers.
Key Points
- Current query routers ignore latency, relying on basic load-balancing methods.
- The new system estimates time-to-first-token (TTFT) for improved routing decisions.
- Joint optimization of latency, accuracy, and cost leads to significant performance gains.
- Experimental results show a 40% improvement in accuracy-cost utility.
- The approach maintains latencies comparable to traditional methods.
DeepSignal Analysis
What happened
A new latency-aware query routing system for language models has been developed, which optimizes query assignments based on latency, accuracy, and cost. This system reportedly improves accuracy-cost utility by up to 40% while maintaining standard latencies. The approach addresses the limitations of existing latency-agnostic routers.
Key evidence
- Current query routers often ignore latency, relying on load-balancing methods like round-robin, which do not consider model accuracy or inference costs.
- The proposed latency estimator simulates autoregressive token batch processing to estimate time-to-first-token (TTFT) for queries.
- Experimental results show that the new routing system achieves a 40% improvement in accuracy-cost utility without increasing latencies compared to standard methods.
Why it matters
Incorporating latency into query routing can significantly enhance the performance of language models, particularly in dynamic workloads where response time is critical. This advancement could lead to more efficient resource utilization and improved user experiences in applications relying on language models. As demand for real-time processing grows, such innovations may become essential for maintaining competitive advantages in AI-driven services.
Paper Resources
📖 Reader Mode
~2 min readAbstract:Modern language query routers improve inference efficiency by assigning each query to a model that balances response quality and monetary cost. However, current query routers are largely latency-agnostic and do not consider the generation latency experienced by queries at model instances. In practice, latency is often controlled by load-balancing policies such as round-robin or join-the-shortest-queue, which do not account for model accuracy or inference cost. Incorporating query latency into routing is challenging as it depends not only on the query's prompt length, but also on the current prefill and decode workload at the model instance and the scheduling and batching policy of the serving framework. We design a lightweight latency estimator that simulates autoregressive token batch processing in the serving framework and estimates the time-to-first-token (TTFT) of queries. We incorporate this latency estimator into a latency-aware router that jointly optimizes latency, accuracy, and cost when assigning queries to model instances. Our experimental results indicate that this joint optimization yields up to 40% improvement in accuracy--cost utility while maintaining the same latencies as standard load-balancing approaches.
| Subjects: | Artificial Intelligence (cs.AI) |
| Cite as: | arXiv:2607.18253 [cs.AI] |
| (or arXiv:2607.18253v1 [cs.AI] for this version) | |
| https://doi.org/10.48550/arXiv.2607.18253 arXiv-issued DOI via DataCite |
Submission history
From: Shivam Patel [view email]
[v1]
Wed, 13 May 2026 20:29:09 UTC (1,247 KB)
— Originally published at arxiv.org
Want this in your inbox every morning?
Daily brief at your local 8am — bilingual EN/中文, free.
More from arXiv cs.AI
See more →HOBA: Hierarchical On-Policy Bidding Agents for Adaptive Online Advertising
HOBA (Hierarchical On-policy Bidding Agents) is a novel hierarchical reinforcement learning framework that enhances online advertising bidding systems by improving adaptability and reducing hyperparameter tuning costs. It utilizes a for hyperparameter inference, a SARSA agent for expert model selection, and a dynamic expert pool for bid execution, achieving a +3.6% increase in target cost during large-scale deployment and outperforming state-of-the-art baselines on AuctionNet.