Routing Without Training: Controllable-Ratio LLM Offloading via Reliability Gating
Quick Answer
CARGO introduces a training-free routing framework for local-cloud collaboration in LLMs, leveraging inference-time agreement for decision-making.
Quick Take
It consistently outperforms training-based routers across various tasks and models, suggesting effective collaboration without additional training. This approach enhances resource efficiency and adaptability in deploying .
Key Points
- CARGO uses prompt-varied sampling to estimate inference-time agreement.
- Employs Bayesian early stopping for efficient uncertainty control.
- Supports arbitrary collaboration ratios with lightweight calibration.
- Outperforms training-free baselines and some supervised routers.
- Demonstrates effectiveness across diverse reasoning and question-answering tasks.
DeepSignal Analysis
What happened
CARGO is a new routing framework that enables local-cloud collaboration for large language models without requiring training. It utilizes inference-time agreement from local models to determine when to execute tasks locally or offload them to cloud models. This approach has shown to outperform traditional training-based routers across various tasks and models.
Key evidence
- CARGO estimates inference-time agreement through prompt-varied sampling, allowing it to make effective routing decisions without prior training.
- The framework consistently outperforms other training-free baselines and, in some cases, even surpasses supervised learned routers across diverse reasoning and question-answering tasks.
- CARGO supports arbitrary target collaboration ratios and employs Bayesian early stopping for efficient uncertainty control, enhancing resource efficiency in deploying large language models.
Why it matters
The introduction of CARGO could significantly reduce the resource demands typically associated with deploying large language models by eliminating the need for additional training. This could lead to more adaptable and efficient systems, particularly in environments with limited resources. The findings suggest that local models can autonomously determine their reliability, which may streamline operations in various applications.
Paper Resources
📖 Reader Mode
~2 min readAbstract:Local-cloud collaboration is a practical way to deploy large language models under resource constraints, but existing methods often rely on trained routers or collaboration-aware finetuning that tie routing behavior to a particular operating regime. In this work, we show that such training may be unnecessary: the local model's own inference-time agreement across sampled responses already provides a strong signal for deciding when to trust local execution and when to offload to a stronger cloud model. We propose CARGO, a training-free routing framework that estimates this agreement through prompt-varied sampling, applies Bayesian early stopping for sample-efficient uncertainty control, and supports arbitrary target collaboration ratios through lightweight deployment-time calibration. Across diverse reasoning and question-answering tasks, multiple local LLM families and scales, and both pretrained and finetuned local models, CARGO consistently outperforms other training-free baselines and in several settings surpasses supervised learned routers. These results suggest that effective and adaptable local-cloud collaboration can emerge directly from the local model's intrinsic response behavior, without requiring an additional trained router.
| Subjects: | Artificial Intelligence (cs.AI) |
| Cite as: | arXiv:2607.20481 [cs.AI] |
| (or arXiv:2607.20481v1 [cs.AI] for this version) | |
| https://doi.org/10.48550/arXiv.2607.20481 arXiv-issued DOI via DataCite (pending registration) |
Submission history
From: Evan Chen [view email]
[v1]
Sat, 30 May 2026 06:42:55 UTC (4,153 KB)
— Originally published at arxiv.org
Want this in your inbox every morning?
Daily brief at your local 8am — bilingual EN/中文, free.
More from arXiv cs.AI
See more →HOBA: Hierarchical On-Policy Bidding Agents for Adaptive Online Advertising
HOBA (Hierarchical On-policy Bidding Agents) is a novel hierarchical reinforcement learning framework that enhances online advertising bidding systems by improving adaptability and reducing hyperparameter tuning costs. It utilizes a for hyperparameter inference, a SARSA agent for expert model selection, and a dynamic expert pool for bid execution, achieving a +3.6% increase in target cost during large-scale deployment and outperforming state-of-the-art baselines on AuctionNet.