When to Rethink: Learning Multi-Perspective Self-Verification for Vision-Language Models
Quick Answer
The paper introduces MOTIVE, a Multi-View Self-Verification framework that enhances the reliability of Vision-Language Models (VLMs) by evaluating answers from multiple perspectives.
Quick Take
Extensive experiments show that MOTIVE outperforms existing self-verification methods, improving decision-making in multimodal reasoning without external judges.
Key Points
- MOTIVE uses multi-view verification to enhance answer reliability in .
- Stronger verifiers yield more reliable judgments, sensitive to prompt choices.
- The framework allows direct return of reliable answers and triggers rethinking for uncertain ones.
- Extensive benchmarks demonstrate MOTIVE's superiority over self-verification baselines.
- Reliable verification reduces unnecessary reasoning turns, improving efficiency.
Paper Resources
📖 Reader Mode
~2 min readAuthors:Ziquan Zhu, Hanruo Zhu, Si-Yuan Lu, Morris Yu-Chao Huang, Yicheng Lin, Wei Han, Tianlong Chen, Mingyuan Wu, Hanchao Yu, Gaojie Jin, Lu Liu, Bo Sun, Tianjin Huang
Abstract:Vision-language models (VLMs) have achieved strong performance in multimodal reasoning, yet they remain prone to generating plausible but incorrect answers. Self-verification offers a practical way to improve answer reliability without relying on external judges, but existing methods typically depend on a single verification criterion or fixed prompt, resulting in incomplete and unstable reliability estimates. We first systematically analyze how verifier capability and prompt design affect verification performance. Our findings show that stronger verifiers provide more reliable judgments, while verification performance is highly sensitive to prompt choice, with no single prompt consistently dominating across tasks. Guided by these findings, we propose \texttt{MOTIVE}, a \textbf{M}ulti-View Self-Verificati\textbf{O}n wi\textbf{T}h Rel\textbf{I}ability-Guided Selecti\textbf{VE} Rethinking framework for reliable multimodal reasoning. \texttt{MOTIVE} evaluates each candidate answer from complementary verification perspectives and learns a correctness-aligned reliability score through correctness-grounded multi-view verification learning. During inference, this score governs an accept-or-rethink decision, allowing reliable answers to be returned directly while uncertain ones trigger history-guided rethinking. Extensive experiments across diverse multimodal benchmarks and VLM backbones demonstrate that \texttt{MOTIVE} consistently outperforms strong self-verification and self-correction baselines. Further results show that reliable verification improves accept-or-rethink decisions and reduces unnecessary reasoning turns, enabling more reliable and efficient self-verification without an external judge.
| Subjects: | Artificial Intelligence (cs.AI) |
| Cite as: | arXiv:2610.07018 [cs.AI] |
| (or arXiv:2610.07018v1 [cs.AI] for this version) | |
| https://doi.org/10.48550/arXiv.2610.07018 arXiv-issued DOI via DataCite (pending registration) |
Submission history
From: Ziquan Zhu [view email]
[v1]
Sun, 4 Oct 2026 15:08:07 UTC (9,818 KB)
— Originally published at arxiv.org
Want this in your inbox every morning?
Daily brief at your local 8am — bilingual EN/中文, free.
More from arXiv cs.AI
See more →HOBA: Hierarchical On-Policy Bidding Agents for Adaptive Online Advertising
HOBA (Hierarchical On-policy Bidding Agents) is a novel hierarchical reinforcement learning framework that enhances online advertising bidding systems by improving adaptability and reducing hyperparameter tuning costs. It utilizes a for hyperparameter inference, a SARSA agent for expert model selection, and a dynamic expert pool for bid execution, achieving a +3.6% increase in target cost during large-scale deployment and outperforming state-of-the-art baselines on AuctionNet.