Region-Aware CLS Token Augmentation for Fine-Grained Image Retrieval
Quick Answer
The study introduces a region-aware CLS token augmentation method for fine-grained image retrieval using DINOv2-reg, enhancing retrieval performance by incorporating localized ROI tokens.
Quick Take
This approach allows for improved discrimination and efficiency in multi-vector retrieval without the need for external bounding boxes, outperforming traditional single-vector methods.
Key Points
- Augments CLS and register tokens with spatial tokens for better semantic representation.
- Automatically captures important regions of interest without external modules.
- Multi-vector retrieval improves performance over DINOv2-reg single-vector baseline.
- Register tokens provide fine-grained details complementing the global CLS token.
- Code available for public access to facilitate further research.
Paper Resources
📖 Reader Mode
~2 min readAbstract:Image retrieval methods often rely on a single global semantic descriptor extracted from an image, e.g., the [CLS] token in vision transformers. However, trying to squeeze all the semantic information of an image into a single descriptor can hurt downstream retrieval performance, especially for fine-grained retrieval tasks. In this work, we augment the semantic tokens in the newer visual transformers, the global [CLS] token and the four register tokens, with a carefully selected collection of spatial tokens, aiming to capture the spatial region representation that characterizes the contents captured in each of the semantic tokens. We leverage the DINOv2-reg model, which includes register tokens that emergently learn object and part-based representations. For each "cue" token ([CLS] and each register token), we find a "buddy" image patch token and extract an N x N patch region to produce a set of localized ROI tokens. Our approach automatically captures important regions of interest without any external bounding boxes or saliency modules, purely by matching semantic tokens with their spatial representation regions. Furthermore, we incorporate these tokens into a multi-vector retrieval framework inspired by ColBERT, enabling fine-grained matching via a per-token alignment mechanism while avoiding the large storage cost of keeping all patch embeddings. Through extensive experiments, we find that (1) register tokens encode useful fine-grained details that can complement the [CLS] token; (2) automatically pooled ROI tokens further improve fine-grained discrimination; and (3) multi-vector retrieval with a small set of tokens improves over a DINOv2-reg single-vector baseline while remaining tractable for large-scale search. The code is available at this https URL.
| Subjects: | Computer Vision and Pattern Recognition (cs.CV); Machine Learning (cs.LG) |
| Cite as: | arXiv:2610.10991 [cs.CV] |
| (or arXiv:2610.10991v1 [cs.CV] for this version) | |
| https://doi.org/10.48550/arXiv.2610.10991 arXiv-issued DOI via DataCite (pending registration) |
Submission history
From: Lucas Pascotti Valem [view email]
[v1]
Wed, 7 Oct 2026 23:24:52 UTC (534 KB)
— Originally published at arxiv.org
Want this in your inbox every morning?
Daily brief at your local 8am — bilingual EN/中文, free.
More from arXiv cs.CV
See more →ProMoE-FL: Prototype-conditioned Mixture of Experts for Multimodal Federated Learning with Missing Modalities
ProMoE-FL introduces a Prototype-conditioned Mixture-of-Experts framework for multimodal federated learning, effectively addressing missing modalities. It outperforms existing methods on four chest X-ray datasets, demonstrating superior feature synthesis capabilities in both homogeneous and heterogeneous settings.