Perception Before Reasoning: Dynamic Latent Reasoning for Video Understanding and Question Answering
Quick Answer
This paper shows that Dynamic Latent Reasoning (DyLaR) enhances video question answering by grounding queries in visual evidence and adaptively deciding when to reason, achieving an accuracy increase from 54.0 to 58.2 on Qwen3-VL-4B while reducing response length from 1,220.7 to 18.5 tokens.
Quick Take
This method outperforms existing baselines across nine benchmarks, demonstrating the effectiveness of grounded perception and adaptive reasoning.
Key Points
- DyLaR grounds questions in perception latents before reasoning.
- Achieved 58.2 average accuracy on Qwen3-VL-4B, up from 54.0.
- Reduced average response length from 1,220.7 to 18.5 tokens.
- Improvements validated across nine video benchmarks.
- Utilizes reinforcement learning for adaptive reasoning decisions.
Paper Resources
📖 Reader Mode
~2 min readAbstract:Video question answering requires models to ground language queries in visual evidence and, when necessary, reason over that evidence across time. Existing methods typically rely on long textual chain-of-thought rationales, even though many questions can be answered as soon as the relevant object, action, or frame is localized. We propose Dynamic Latent Reasoning (DyLaR), which first grounds a question in a short block of perception latents (continuous hidden states that encode query-relevant visual evidence), and then adaptively decides whether to append reasoning latents (continuous thoughts that reason over this evidence in latent space) before answering. DyLaR learns this behavior by grounding perception latents in verified visual evidence and distilling verified rationales into reasoning latents, followed by reinforcement learning that further refines when to reason. Across nine video benchmarks and four multimodal language model backbones, DyLaR improves average accuracy over same-backbone baselines while generating fewer than 20 tokens per query. On Qwen3-VL-4B, for example, DyLaR improves average accuracy over Qwen3-VL-4B-Thinking from 54.0 to 58.2 while reducing response length from 1,220.7 to 18.5 tokens per query. Ablations further show that grounded perception latents, rationale-supervised reasoning latents, and adaptive routing each improve accuracy.
| Subjects: | Computer Vision and Pattern Recognition (cs.CV); Artificial Intelligence (cs.AI) |
| Cite as: | arXiv:2608.04124 [cs.CV] |
| (or arXiv:2608.04124v1 [cs.CV] for this version) | |
| https://doi.org/10.48550/arXiv.2608.04124 arXiv-issued DOI via DataCite (pending registration) |
Submission history
From: Haotian Xia [view email]
[v1]
Tue, 4 Aug 2026 18:23:17 UTC (3,914 KB)
— Originally published at arxiv.org
Want this in your inbox every morning?
Daily brief at your local 8am — bilingual EN/中文, free.
More from arXiv cs.CV
See more →ProMoE-FL: Prototype-conditioned Mixture of Experts for Multimodal Federated Learning with Missing Modalities
ProMoE-FL introduces a Prototype-conditioned Mixture-of-Experts framework for multimodal federated learning, effectively addressing missing modalities. It outperforms existing methods on four chest X-ray datasets, demonstrating superior feature synthesis capabilities in both homogeneous and heterogeneous settings.