Ask Twice, Look Twice: Prompt Echoing Resolves the Question-First Paradox in Vision-Language Models
Quick Answer
The study reveals that prompt echoing in vision-language models (VLMs) resolves the question-first paradox, enhancing performance by up to 19 accuracy points on benchmarks like NaturalBench and VQAv2.
Quick Take
By restating questions before and after images, models better align perception with query relevance, improving answer accuracy without additional training or architecture changes.
Key Points
- Question-first prompting underperforms compared to image-first in .
- Prompt echoing improves accuracy by up to 19 points on various benchmarks.
- The method requires no training or architecture changes.
- Echoing enhances perception alignment with question relevance.
- The approach is supported by findings on human comprehension of adjunct questions.
DeepSignal Analysis
What happened
The study investigates the positioning of questions in vision-language models (VLMs) and introduces prompt echoing as a solution to the question-first paradox. This method enhances model performance on benchmarks, achieving accuracy improvements of up to 19 points without requiring additional training or changes to model architecture.
Key evidence
- The research identifies a question-first paradox where placing questions before images leads to lower performance in VLMs compared to image-first prompting.
- Prompt echoing involves restating questions before and after images, which helps align perception with query relevance, improving answer accuracy.
- The study reports that echoed prompts can surpass the best single-pass ordering on benchmarks like NaturalBench and VQAv2, achieving up to 19 accuracy points improvement.
Why it matters
This research highlights a significant limitation in current VLMs regarding how questions are integrated into prompts. By addressing the question-first paradox, the findings suggest a straightforward method to enhance model performance without the need for complex retraining or architectural adjustments. This could lead to more effective applications of VLMs in various domains, including visual question answering.
Paper Resources
📖 Reader Mode
~2 min readAbstract:Where should the question go in a vision-language model (VLM) prompt: before the image or after it? Intuition says before: knowing what is asked should tell the model where to look. Yet across visual question answering benchmarks, question-first prompting consistently underperforms the image-first ordering recommended for frontier VLMs, a phenomenon we term the question-first paradox. We trace the paradox to a conflict between two stages of VLM computation. Logit-lens and attention probes show the intuition is half right: a question placed before the image genuinely steers perception, moving image patch representations toward question-relevant concepts. The failure lies downstream. Stranded behind hundreds of image tokens, the question is barely attended by the answer token, which instead commits to image-driven (often wrong) answers; a causal attention knockout confirms that the answer reads the question only when the question follows the image. The diagnosis yields a training-free fix: question echoing, restating the question on both sides of the image so that one copy steers perception while the other is read out at answer time. The same division of labor appears in a fifty-year-old finding on human ``adjunct questions'', where repeating a question before and after a passage aids comprehension more than either position alone. Echoing the image as well brings further gains, restoring the whole-image view a causal decoder otherwise loses. The paradox holds across five open VLMs, costing up to 17.5 group-accuracy points. Echoed prompts close it and surpass the best single-pass ordering on NaturalBench, POPE, Winoground, and open-ended VQAv2, by up to 19 Winoground group-accuracy points, with no training, fine-tuning, or architecture change. The paradox reveals a trade-off between steering perception and preserving question access; echoing resolves it through prompt design alone.
| Subjects: | Computer Vision and Pattern Recognition (cs.CV); Artificial Intelligence (cs.AI); Machine Learning (cs.LG); Image and Video Processing (eess.IV) |
| Cite as: | arXiv:2607.15565 [cs.CV] |
| (or arXiv:2607.15565v1 [cs.CV] for this version) | |
| https://doi.org/10.48550/arXiv.2607.15565 arXiv-issued DOI via DataCite (pending registration) |
Submission history
From: Gautam Gare [view email]
[v1]
Fri, 17 Jul 2026 02:16:46 UTC (6,707 KB)
— Originally published at arxiv.org
Want this in your inbox every morning?
Daily brief at your local 8am — bilingual EN/中文, free.
More from arXiv cs.CV
See more →ProMoE-FL: Prototype-conditioned Mixture of Experts for Multimodal Federated Learning with Missing Modalities
ProMoE-FL introduces a Prototype-conditioned Mixture-of-Experts framework for multimodal federated learning, effectively addressing missing modalities. It outperforms existing methods on four chest X-ray datasets, demonstrating superior feature synthesis capabilities in both homogeneous and heterogeneous settings.