IMPRINT: Image-Conditioned Query Enrichment for Long-Tail Object Goal Navigation
Quick Answer
IMPRINT introduces a framework for enhancing zero-shot Object Goal Navigation by enriching text queries with web-sourced images, improving grounding in semantic maps without requiring training.
Quick Take
The new HSSD-rare benchmark demonstrates significant gains in navigation performance, particularly for long-tail object categories, highlighting the importance of downstream detection quality.
Key Points
- IMPRINT enhances Object Goal Navigation using image-conditioned queries for better localization.
- No training or modification of navigation policy is required for implementation.
- HSSD-rare benchmark features semantically specific subcategories for long-tail evaluation.
- Image-conditioned queries consistently improve object grounding and navigation performance.
- Downstream detection quality is critical for translating localization gains into navigation success.
Paper Resources
📖 Reader Mode
~2 min readAbstract:Embodied AI increasingly relies on queryable semantic maps built from pre-trained vision-language models to enable zero-shot Object Goal Navigation (ObjectNav). However, existing approaches typically depend on text-only queries, which become less reliable as semantic specificity increases toward fine-grained object categories. We introduce IMPRINT, a zero-shot plug-and-play framework that enriches textual object queries with web-sourced images to improve grounding in queryable maps. Retrieved images are encoded using a vision-language model, matched against the semantic map to produce similarity maps, and aggregated to yield context-aware localization. Notably, this requires no training or modification of the underlying navigation policy. To explicitly evaluate long-tail behavior, we present HSSD-rare, a new ObjectNav benchmark built on Habitat Synthetic Scenes and featuring semantically specific subcategories. Across both OVON and HSSD-rare, image-conditioned queries consistently improve object grounding and yield end-to-end navigation gains. Further analysis reveals that translating localization gains to navigation performance depends critically on downstream detection quality, highlighting a key systems bottleneck in long-tail embodied navigation.
| Comments: | Accepted at the 2026 IEEE/RSJ International Conference on Intelligent Robots and Systems (IROS). Project page: this https URL |
| Subjects: | Computer Vision and Pattern Recognition (cs.CV) |
| Cite as: | arXiv:2607.25106 [cs.CV] |
| (or arXiv:2607.25106v1 [cs.CV] for this version) | |
| https://doi.org/10.48550/arXiv.2607.25106 arXiv-issued DOI via DataCite (pending registration) |
Submission history
From: Jelin Raphael Akkara [view email]
[v1]
Mon, 27 Jul 2026 22:09:32 UTC (4,110 KB)
— Originally published at arxiv.org
Want this in your inbox every morning?
Daily brief at your local 8am — bilingual EN/中文, free.
More from arXiv cs.CV
See more →ProMoE-FL: Prototype-conditioned Mixture of Experts for Multimodal Federated Learning with Missing Modalities
ProMoE-FL introduces a Prototype-conditioned Mixture-of-Experts framework for multimodal federated learning, effectively addressing missing modalities. It outperforms existing methods on four chest X-ray datasets, demonstrating superior feature synthesis capabilities in both homogeneous and heterogeneous settings.