What Makes Linguistic Representations Good Models of High-Level Visual Perception in the Human Brain?
Quick Answer
This study reveals that machine-generated captions using language models (LMs) significantly predict human brain responses to images, outperforming human-annotated captions.
Quick Take
Text embedders consistently excel over autoregressive LMs, with peak predictivity at intermediate network depths, highlighting the importance of caption content and model choice in visual perception research.
Key Points
- Machine-generated captions show higher brain predictivity than human-annotated captions.
- Text embedders outperform autoregressive LMs in predicting brain responses.
- Peak predictivity occurs at intermediate network depths in LMs.
- Study utilized six caption types across a common image set.
- Findings support using caption embeddings for high-level visual perception analysis.
DeepSignal Analysis
What happened
The study investigates how different types of image captions, generated by various language models (LMs), predict human brain responses to images. It finds that machine-generated captions often outperform human-annotated ones, with text embedders showing consistent superiority over autoregressive LMs. Additionally, brain predictivity peaks at intermediate network depths.
Key evidence
- Machine-generated captions significantly predict human brain responses to images, often surpassing human-annotated captions used in previous studies.
- Text embedders consistently outperform autoregressive LMs across different caption types, indicating a preference for certain model architectures in visual perception tasks.
- Brain predictivity and behavioral alignment peak at intermediate network depths, suggesting a critical point for the emergence of syntactic and semantic structures in LMs.
Why it matters
Understanding how linguistic representations relate to visual perception can enhance models used in AI and cognitive science. The findings suggest that the choice of language model and the content of captions are crucial for accurately modeling human brain responses. This could lead to improved AI systems that better mimic human cognitive processes, particularly in interpreting visual information.
Paper Resources
📖 Reader Mode
~2 min readAbstract:Image descriptions represented with language models (LMs) predict human brain responses to naturalistic images in high-level visual regions, but the factors driving this predictivity remain unclear. To investigate this, we systematically studied how images are described and which language models are used to embed those descriptions. For a common set of images, we considered six caption types -- including human-annotated and multiple machine-generated captions -- differing along several dimensions. Each caption was represented with five LMs, spanning autoregressive LMs trained to predict upcoming words and text embedders, i.e., LMs fine-tuned on semantic tasks requiring sentence/document-level representations. Machine-generated captions yielded significant brain predictivity and alignment, often surpassing human-annotated captions used in previous work. Across caption types, text embedders consistently outperformed autoregressive LMs, a pattern replicated when measuring behavioural alignment with image-similarity judgments. Analyses of caption representations from different model layers further revealed that both brain predictivity and behavioural alignment peak at intermediate network depth, shortly after a point thought to mark the emergence of syntactic and semantic structure. Altogether, our results demonstrate that both the content of image captions and the LM used to represent them influence brain- and behaviour-modelling performance, establishing caption embeddings as a useful tool for studying high-level visual perception.
| Comments: | Supplementary Materials in the public GitHub repository referenced in the paper |
| Subjects: | Computer Vision and Pattern Recognition (cs.CV); Artificial Intelligence (cs.AI) |
| Cite as: | arXiv:2607.16214 [cs.CV] |
| (or arXiv:2607.16214v1 [cs.CV] for this version) | |
| https://doi.org/10.48550/arXiv.2607.16214 arXiv-issued DOI via DataCite |
Submission history
From: Anna Bavaresco [view email]
[v1]
Fri, 22 May 2026 16:43:14 UTC (1,476 KB)
— Originally published at arxiv.org
Want this in your inbox every morning?
Daily brief at your local 8am — bilingual EN/中文, free.
More from arXiv cs.CV
See more →ProMoE-FL: Prototype-conditioned Mixture of Experts for Multimodal Federated Learning with Missing Modalities
ProMoE-FL introduces a Prototype-conditioned Mixture-of-Experts framework for multimodal federated learning, effectively addressing missing modalities. It outperforms existing methods on four chest X-ray datasets, demonstrating superior feature synthesis capabilities in both homogeneous and heterogeneous settings.