Metonymic Circuits for Abstract Concept Grounding in Vision Transformers
Quick Answer
This study proposes a metonymic grounding mechanism in Vision Transformers, where abstract concepts like 'angry' are linked to concrete anchors such as 'fire'.
Quick Take
By utilizing Transcoders on CLIP and DINO encoders, the research identifies structured circuits that facilitate abstract concept recognition, validated through causal interventions on a curated icon dataset.
Key Points
- Metonymic grounding links abstract concepts to concrete anchors in Vision Transformers.
- Transcoders applied on CLIP and DINO recover features associated with concrete concepts.
- Structured circuits show perceptual primitives dominate early layers of recognition.
- Causal interventions validate the functional role of metonymic intermediates.
- Distinct routes are observed for images with rendered text versus abstract targets.
Paper Resources
📖 Reader Mode
~2 min readAbstract:We study how Vision Transformers ground abstract concepts (e.g., angry) when training data provide limited direct referential evidence. We hypothesize a metonymic grounding mechanism in which abstract predictions are driven by concrete, interpretable anchor concepts (e.g., fire) that bridge visual signals to abstract semantics. By applying Transcoders on CLIP and DINO vision encoders, we recover intermediate features that can be associated with semantic labels for more concrete concepts, and trace their contributions in circuits underlying abstract concept recognition. Experiments on a carefully curated icon dataset reveal structured metonymic circuits, in which perceptual primitives dominate early layers and object-like anchors precede abstract targets. Images containing rendered text instead recruit a distinct perceptual-to-textual route. Causal interventions further validate that metonymic intermediates are functionally involved in grounding abstract concepts.
| Comments: | EMNLP 2026 Main. Project Website: this https URL |
| Subjects: | Artificial Intelligence (cs.AI) |
| Cite as: | arXiv:2610.06928 [cs.AI] |
| (or arXiv:2610.06928v1 [cs.AI] for this version) | |
| https://doi.org/10.48550/arXiv.2610.06928 arXiv-issued DOI via DataCite |
Submission history
From: Jing Ding [view email]
[v1]
Sat, 3 Oct 2026 00:53:32 UTC (3,894 KB)
— Originally published at arxiv.org
Want this in your inbox every morning?
Daily brief at your local 8am — bilingual EN/中文, free.
More from arXiv cs.AI
See more →HOBA: Hierarchical On-Policy Bidding Agents for Adaptive Online Advertising
HOBA (Hierarchical On-policy Bidding Agents) is a novel hierarchical reinforcement learning framework that enhances online advertising bidding systems by improving adaptability and reducing hyperparameter tuning costs. It utilizes a for hyperparameter inference, a SARSA agent for expert model selection, and a dynamic expert pool for bid execution, achieving a +3.6% increase in target cost during large-scale deployment and outperforming state-of-the-art baselines on AuctionNet.