CofactVLA: Deconfounding Vision-Language-Action Models via Counterfactual Intervention
Quick Answer
CofactVLA introduces a novel causal intervention framework to address the vision-override issue in Vision-Language-Action models, achieving state-of-the-art results in simulations and a 52.3% success rate improvement in real-world robotic tasks under out-of-distribution scenarios.
Key Points
- CofactVLA utilizes a Dual-path Deconfounding Graph for action generation.
- Action-Level Orthogonal Projection Guidance isolates semantic intent from visual biases.
- Feature-Level Counterfactual Covariance Reduction suppresses dominant visual shortcuts.
- Extensive experiments show CofactVLA sets new benchmarks across various simulations.
- Real-world tests demonstrate significant improvements in generalization capabilities.
DeepSignal Analysis
What happened
CofactVLA is a new framework designed to address the vision-override issue in Vision-Language-Action models. It utilizes a Dual-path Deconfounding Graph to improve action generation by isolating visual confounders. The framework has shown a 52.3% improvement in success rates for robotic tasks in out-of-distribution scenarios.
Key evidence
- CofactVLA introduces a Dual-path Deconfounding Graph to formalize action generation, aiming to reduce bias from visual confounders.
- The framework employs Action-Level Orthogonal Projection Guidance and Feature-Level Counterfactual Covariance Reduction to enhance the model's performance.
- In real-world robotic tasks, CofactVLA achieved a 52.3% absolute success rate improvement under out-of-distribution conditions.
Why it matters
The vision-override phenomenon poses significant challenges for Vision-Language-Action models, often leading to reliance on misleading visual cues rather than accurate language instructions. By addressing this issue, CofactVLA could enhance the reliability and effectiveness of robotic manipulation tasks, particularly in unpredictable environments. This advancement may have broader implications for the integration of AI in real-world applications where understanding context is crucial.
Paper Resources
📖 Reader Mode
~2 min readAbstract:Vision-Language-Action (VLA) models have driven significant progress in robotic manipulation, yet they fundamentally struggle with the vision-override phenomenon. Driven by the severe modality imbalance between dense visual streams and sparse linguistic instructions, VLAs frequently fall prey to causal confusion. Instead of treating language as the primary causal driver, the policy entirely bypasses the original instruction by overfitting to spurious visual confounders, such as prominent objects or familiar layouts. To systematically alleviate this bias, we formalize the process of action generation as a Dual-path Deconfounding Graph (DDG) and propose CofactVLA, a novel causal intervention framework. By dynamically constructing a language-masked counterfactual branch within a single forward pass, CofactVLA isolates and neutralizes visual confounders through two synergistic mechanisms. First, Action-Level Orthogonal Projection Guidance (OPG) geometrically projects the factual velocity field away from the counterfactual visual bias during continuous flow matching, extracting the pure semantic intent. Second, Feature-Level Counterfactual Covariance Reduction (CCR) mathematically deconfounds latent representations by penalizing the positive eigenspace of the covariance difference, explicitly suppressing dominant visual shortcuts while preserving the causal language intent. Extensive experiments demonstrate that CofactVLA establishes a new state-of-the-art across diverse simulation benchmarks. Beyond simulation, real-world robot experiments demonstrate the causal efficacy of our method in bridging the generalization gap, yielding a 52.3\% absolute success rate gain under out-of-distribution scenarios.
| Subjects: | Computer Vision and Pattern Recognition (cs.CV) |
| Cite as: | arXiv:2608.04396 [cs.CV] |
| (or arXiv:2608.04396v1 [cs.CV] for this version) | |
| https://doi.org/10.48550/arXiv.2608.04396 arXiv-issued DOI via DataCite (pending registration) |
Submission history
From: Yan Zhang [view email]
[v1]
Wed, 5 Aug 2026 02:58:41 UTC (13,007 KB)
— Originally published at arxiv.org
Want this in your inbox every morning?
Daily brief at your local 8am — bilingual EN/中文, free.
More from arXiv cs.CV
See more →ProMoE-FL: Prototype-conditioned Mixture of Experts for Multimodal Federated Learning with Missing Modalities
ProMoE-FL introduces a Prototype-conditioned Mixture-of-Experts framework for multimodal federated learning, effectively addressing missing modalities. It outperforms existing methods on four chest X-ray datasets, demonstrating superior feature synthesis capabilities in both homogeneous and heterogeneous settings.