CARE: Certifying Acceleration for Vision-Language-Action Inference
Quick Answer
CARE introduces a certified approach for accelerating vision-language-action inference, achieving 9.0–10.8x speedups while ensuring at least 85.8% of reference-solved episodes are preserved.
Quick Take
It utilizes paired rollouts to manage acceleration-induced failures, outperforming traditional selectors that exceed budget in 75% of trials.
Key Points
- CARE certifies acceleration methods while maintaining a user-specified failure risk budget.
- Achieves 9.0–10.8x speedups on four LIBERO suites with OpenVLA-OFT.
- Ensures 85.8% of reference episodes are preserved at 95% confidence.
- Uses 78.9% fewer rollouts than exhaustive evaluation under tight budgets.
- Generalizes to flow-step reduction for various agents including Qwen3.5-9B.
Paper Resources
📖 Reader Mode
~2 min readAbstract:While vision-language-action (VLA) models have advanced rapidly, running them at every control step remains expensive. Prior work accelerates VLA inference using techniques like action chunking and visual-token pruning, typically evaluating based on latency and average task success. However, acceleration may discard information and break tasks the original policy would solve, a risk hidden by average metrics. Measuring these failures is challenging because action deviations compound over closed-loop trajectories, meaning task failure is only observable across full episodes. We therefore define an acceleration-induced failure via paired rollouts from identical initial conditions, tracking when the reference succeeds but the accelerated policy fails. To manage this, we introduce CARE, an approach for certified accelerator selection. CARE uses paired rollouts on a calibration set to provide finite-sample guarantees that acceleration-induced failure risk stays below a user-specified budget. It deploys the fastest certified candidate, falling back to the reference if none qualify. By relying only on terminal outcomes and measured compute, CARE applies unchanged across diverse acceleration mechanisms, while sequential testing and failure-triggered reference rollouts keep certification affordable. On four LIBERO suites with OpenVLA-OFT, CARE certifies $9.0$--$10.8\times$ speedups while guaranteeing (at $95\%$ confidence) that at least $85.8\%$ of reference-solved episodes are preserved. Under tight budgets, selectors without guarantees exceed the budget in up to $75\%$ of trials, whereas CARE stays within budget and its sequential form uses $78.9\%$ fewer rollouts than exhaustive evaluation. CARE further generalizes to flow-step reduction for $\pi_{0.5}$, and to Qwen3.5-9B and Llama-3.1-8B agents in Crafter.
| Subjects: | Computation and Language (cs.CL); Machine Learning (cs.LG) |
| Cite as: | arXiv:2610.08917 [cs.CL] |
| (or arXiv:2610.08917v1 [cs.CL] for this version) | |
| https://doi.org/10.48550/arXiv.2610.08917 arXiv-issued DOI via DataCite (pending registration) |
Submission history
From: Rui Liu [view email]
[v1]
Tue, 6 Oct 2026 18:00:05 UTC (2,105 KB)
— Originally published at arxiv.org
Want this in your inbox every morning?
Daily brief at your local 8am — bilingual EN/中文, free.
More from arXiv cs.CL
See more →TriAgent: Divergence-Aware Committees for Cost-Efficient Financial Sentiment Analysis
TriAgent introduces a cost-efficient multi-agent system for financial sentiment analysis, combining VADER, FinBERT, and Qwen2.5. It achieves an F1 score of ~0.87 with significant savings of $9.3M/year at a 10M-user scale compared to GPT-4o-mini, while also detecting hallucinations with an AUC of 0.90.