Rethinking Feature Reliance Evaluation with Semantically Matched Suppression
Quick Answer
The study introduces a semantically matched evaluation framework revealing that ImageNet-trained CNNs exhibit greater reliance on texture than shape, challenging prior views of CNN bias.
Quick Take
Additionally, Vision Transformers outperform CNNs in accuracy under both suppression types, suggesting their representations align better with human visual processing.
Key Points
- CNNs show stronger performance degradation under texture suppression compared to shape suppression.
- Vision Transformers maintain higher accuracy than CNNs during both shape and texture suppression.
- Semantic comparability is crucial for interpreting feature reliance in suppression experiments.
- Brain encoding indicates ViT representations are less affected by suppression than CNNs.
- Findings suggest ViTs may better align with human visual cortex representations.
DeepSignal Analysis
What happened
The study presents a new evaluation framework that assesses feature reliance in visual recognition models, specifically focusing on shape and texture. It finds that ImageNet-trained CNNs are more reliant on texture than previously thought, while Vision Transformers demonstrate superior accuracy under both suppression types.
Key evidence
- The authors introduce a semantically matched evaluation framework that compares shape and texture suppression, revealing that CNNs show greater reliance on texture.
- Under the new framework, ImageNet-trained CNNs experience more significant performance drops with texture suppression compared to shape suppression.
- Vision Transformers outperform CNNs in accuracy during both shape and texture suppression, indicating a potential alignment with human visual processing.
Why it matters
This research challenges the prevailing notion that CNNs are primarily biased towards texture by demonstrating a more nuanced understanding of feature reliance. The findings suggest that the robustness of Vision Transformers may be linked to their representation being more compatible with human visual processing, which could influence future model design and evaluation methods.
Paper Resources
📖 Reader Mode
~2 min readAuthors:Ning Jiang (1 and 2), Tianyi Luo (4), Zhengyong Huang (1 and 2), Yao Sui (1 and 2 and 3) ((1) Institute of Medical Technology, Peking University Health Science Center, Beijing, China, (2) National Institute of Health Data Science, Peking University, Beijing, China,(3) Institute for Artificial Intelligence, Peking University, Beijing, China,(4) School of Computer Science and Engineering, Sun Yat-sen University, Guangzhou, China)
Abstract:Understanding whether visual recognition models rely on shape, texture, or color is central to interpreting their behavior. Prior cue-conflict studies have strongly influenced the view that CNNs are texture-biased, yet such tests measure cue preference under artificial conflicts rather than feature reliance during natural recognition. We revisit this question through controlled feature suppression and show that performance drops are difficult to interpret unless different suppression operations impose comparable category-level damage. We introduce a semantically matched evaluation framework that compares shape and texture suppression at matched levels of category separability loss. Under this framework, ImageNet-trained CNNs show stronger degradation under texture suppression than under shape suppression, revealing greater texture reliance than suggested by unmatched suppression analyses. Extending the comparison across architectures, we find that Vision Transformers retain higher accuracy than CNNs under both shape and texture suppression. Brain encoding further shows that ViT representations exhibit smaller suppression-induced decreases in neural prediction performance under the tested suppression settings. These findings indicate that semantic comparability is essential for interpreting feature reliance from suppression experiments, and suggest that the robustness advantage of ViTs may be related to representations more compatible with human visual cortex.
| Subjects: | Computer Vision and Pattern Recognition (cs.CV) |
| Cite as: | arXiv:2607.16298 [cs.CV] |
| (or arXiv:2607.16298v1 [cs.CV] for this version) | |
| https://doi.org/10.48550/arXiv.2607.16298 arXiv-issued DOI via DataCite |
Submission history
From: Tianyi Luo [view email]
[v1]
Mon, 13 Jul 2026 13:02:07 UTC (24,538 KB)
— Originally published at arxiv.org
Want this in your inbox every morning?
Daily brief at your local 8am — bilingual EN/中文, free.
More from arXiv cs.CV
See more →ProMoE-FL: Prototype-conditioned Mixture of Experts for Multimodal Federated Learning with Missing Modalities
ProMoE-FL introduces a Prototype-conditioned Mixture-of-Experts framework for multimodal federated learning, effectively addressing missing modalities. It outperforms existing methods on four chest X-ray datasets, demonstrating superior feature synthesis capabilities in both homogeneous and heterogeneous settings.