Criterion-Conditional In-Context Learning: Evaluating Criterion-Shift Adaptation in Vision-Language Models
Quick Answer
The study introduces Criterion-Conditional In-Context Learning (CC-ICL) for vision-language models, enabling them to adapt to shifting decision criteria without altering task semantics.
Quick Take
Experiments on the CC-Bench benchmark reveal that most models struggle with criterion alignment, but a multi-criterion training approach significantly enhances adaptability, allowing 7B-scale models to outperform proprietary counterparts while maintaining multimodal performance.
Key Points
- CC-ICL allows models to infer latent criteria from context for better decision-making.
- CC-Bench benchmark evaluates models' performance under shifting decision criteria.
- Most models show rigid boundary bias, struggling with criterion alignment.
- Multi-criterion training significantly improves adaptability and criterion sensitivity.
- 7B-scale models can outperform proprietary models without losing multimodal capabilities.
Paper Resources
📖 Reader Mode
~2 min readAbstract:Vision-language models can perform new tasks without parameter updates through in-context learning (ICL), whose core mechanism is utilizing the support set for task induction. In the standard ICL setting, once the task is induced, its decision criterion remains fixed. However, in real-world applications, many tasks exhibit a stable high-level intent, while their decision criteria shift according to specific requirements. Thus, we introduce a new setting, denoted as Criterion-Conditional In-Context Learning (CC-ICL), where models must infer the latent criterion from context and adjust predictions accordingly under fixed task semantics. To evaluate this capability, we propose two complementary metrics, Criterion Invariance and Criterion Sensitivity, capturing the model's robustness and adaptability under criterion shifts. We further construct CC-Bench, a multi-domain benchmark that supports evaluation under the CC-ICL setting. By employing a dual-level data hierarchy, CC-Bench enables legitimate ground-truth variation conditioned on the active criterion even when the task remains fixed. Experiments on CC-Bench reveal that most models exhibit a rigid boundary bias, struggling to align their decisions with the latent criterion. We also find that even a simple multi-criterion training strategy can significantly reduce this bias, improving Criterion Sensitivity and enabling 7B-scale models to surpass proprietary models without degrading general multimodal performance.
| Comments: | Accepted by ICML 2026 |
| Subjects: | Computer Vision and Pattern Recognition (cs.CV); Artificial Intelligence (cs.AI) |
| Cite as: | arXiv:2607.02575 [cs.CV] |
| (or arXiv:2607.02575v1 [cs.CV] for this version) | |
| https://doi.org/10.48550/arXiv.2607.02575 arXiv-issued DOI via DataCite |
Submission history
From: Ruilin Yang [view email]
[v1]
Tue, 30 Jun 2026 17:22:20 UTC (10,771 KB)
— Originally published at arxiv.org
Want this in your inbox every morning?
Daily brief at your local 8am — bilingual EN/中文, free.
More from arXiv cs.CV
See more →ProMoE-FL: Prototype-conditioned Mixture of Experts for Multimodal Federated Learning with Missing Modalities
ProMoE-FL introduces a Prototype-conditioned Mixture-of-Experts framework for multimodal federated learning, effectively addressing missing modalities. It outperforms existing methods on four chest X-ray datasets, demonstrating superior feature synthesis capabilities in both homogeneous and heterogeneous settings.