RACER: Reflective Agent Coupling Query Interpretation and Tool-Based Retrieval for Frame Selection in Long Video Understanding
Quick Answer
RACER introduces a training-free framework for long video frame selection, addressing the Query Comprehension and Interpretation-Selection Gaps.
Quick Take
By utilizing a lightweight Vid- for query reformulation and an embedding model for evidence localization, RACER enhances frame selection effectiveness, demonstrating improved performance across benchmarks even with limited-capability components.
Key Points
- RACER decomposes frame selection into query interpretation and evidence localization.
- The framework mitigates Query Comprehension and Interpretation-Selection Gaps.
- It employs a lightweight Vid-LLM for sub-query reformulation.
- RACER shows consistent improvements in long video understanding across multiple benchmarks.
- Effective frame selection is achieved even with limited-capability components.
Paper Resources
📖 Reader Mode
~2 min readAbstract:Video large language models (Vid-LLMs) excel at diverse video-language tasks by reasoning over selected frames. However, frame selection for long videos remains challenging, as it requires retrieving relevant frames distributed across segments from a large candidate pool given complex queries. This paper investigates dominant approaches to long-video frame selection from a task-decomposition perspective, identifying two key challenges: the Query Comprehension Gap in similarity-based methods and the Interpretation--Selection Gap in judgment-based methods. To address them, we propose RACER, a training-free reflective agentic framework that decomposes long-video frame selection into query interpretation driven by a lightweight Vid-LLM and evidence localization supported by an embedding model serving as a retrieval tool. Specifically, the Vid-LLM is responsible solely for reformulating the complex query into sub-queries that make implicit information requirements explicit, mitigating the Query Comprehension Gap. Meanwhile, the retrieval tool leverages these sub-queries to localize relevant evidence, relieving the Vid-LLM of direct frame selection and thus addressing the Interpretation--Selection Gap. Finally, the retrieved frames are fed back to the Vid-LLM for sub-query refinement, forming a reflection loop that iteratively improves query interpretation and frame selection. Experiments across multiple benchmarks show that RACER consistently improves long video understanding. Notably, RACER achieves effective frame selection even with limited-capability components, demonstrating that agentic integration enables these components to enhance more capable Vid-LLMs.
| Subjects: | Computer Vision and Pattern Recognition (cs.CV); Artificial Intelligence (cs.AI) |
| Cite as: | arXiv:2610.08954 [cs.CV] |
| (or arXiv:2610.08954v1 [cs.CV] for this version) | |
| https://doi.org/10.48550/arXiv.2610.08954 arXiv-issued DOI via DataCite (pending registration) |
Submission history
From: Yiyang Huang [view email]
[v1]
Tue, 6 Oct 2026 18:21:13 UTC (18,646 KB)
— Originally published at arxiv.org
Want this in your inbox every morning?
Daily brief at your local 8am — bilingual EN/中文, free.
More from arXiv cs.CV
See more →ProMoE-FL: Prototype-conditioned Mixture of Experts for Multimodal Federated Learning with Missing Modalities
ProMoE-FL introduces a Prototype-conditioned Mixture-of-Experts framework for multimodal federated learning, effectively addressing missing modalities. It outperforms existing methods on four chest X-ray datasets, demonstrating superior feature synthesis capabilities in both homogeneous and heterogeneous settings.