Less from More: Reinforcing Sparse Video Reasoning from Dense References
Quick Answer
SAVER is a novel dense-to-sparse post-training framework that enhances sparse video reasoning using dense video references, achieving superior performance on temporal grounding and video question-answering benchmarks with only 1,250 training examples.
Quick Take
It matches or exceeds dense-frame Qwen3.5 baselines while utilizing significantly fewer frames, demonstrating effective evidence localization for broader video understanding tasks.
Key Points
- SAVER optimizes sparse video predictions with grounding rewards and reference rewards.
- Trained on only 1,250 temporal grounding examples without video question answering annotations.
- Improves performance across three temporal grounding benchmarks and six video QA benchmarks.
- Matches or surpasses dense-frame Qwen3.5 baselines with fewer frames.
- Demonstrates effective evidence localization for broader video understanding tasks.
Paper Resources
📖 Reader Mode
~2 min readAbstract:Video-language models commonly assume that more temporal observations lead to more reliable reasoning. We question this assumption and argue that the key challenge is not merely processing more video frames efficiently, but learning to reason reliably under limited temporal evidence. We propose SAVER, a dense-to-sparse post-training framework that uses dense video views as training-time references for sparse-frame inference. During reinforcement post-training, paired dense and sparse views are optimized with grounding rewards and a reliability-gated reference reward, encouraging sparse view predictions to preserve task-relevant temporal evidence. Notably, SAVER is trained only on 1,250 randomly sampled temporal grounding examples, without using any video question answering annotations. Across three temporal grounding benchmarks and six video question-answering benchmarks, SAVER consistently improves performance across frame budgets. In particular, SAVER can match or surpass dense-frame Qwen3.5 baselines while using substantially fewer frames. These results show that temporal grounding can serve as an effective evidence-localization proxy for learning sparse video reasoning that transfers to broader video understanding tasks.
| Subjects: | Computer Vision and Pattern Recognition (cs.CV) |
| Cite as: | arXiv:2610.10893 [cs.CV] |
| (or arXiv:2610.10893v1 [cs.CV] for this version) | |
| https://doi.org/10.48550/arXiv.2610.10893 arXiv-issued DOI via DataCite (pending registration) |
Submission history
From: Wenfang Sun [view email]
[v1]
Wed, 7 Oct 2026 20:51:59 UTC (30,525 KB)
— Originally published at arxiv.org
Want this in your inbox every morning?
Daily brief at your local 8am — bilingual EN/中文, free.
More from arXiv cs.CV
See more →ProMoE-FL: Prototype-conditioned Mixture of Experts for Multimodal Federated Learning with Missing Modalities
ProMoE-FL introduces a Prototype-conditioned Mixture-of-Experts framework for multimodal federated learning, effectively addressing missing modalities. It outperforms existing methods on four chest X-ray datasets, demonstrating superior feature synthesis capabilities in both homogeneous and heterogeneous settings.