CROP: Expert-Aligned Image Cropping via Compositional Reasoning and Optimizing Preference

arXiv cs.CV·Zhitong Dong, Chao Li, Jie Yu, Hao Chen

3d ago

·~2 min·5/14/2026·en·1

Quick Take

CROP reformulates aesthetic image cropping as a multimodal reasoning task to align with expert preferences.

Key Points

Introduces Compositional Reasoning and Optimizing Preference method.
Enhances cropping decisions through expert alignment.
Demonstrates superior performance across multiple datasets.

📖 Reader Mode

~2 min read

[Submitted on 9 May 2026]

View PDF HTML (experimental)

Abstract:Aesthetic image cropping aims to enhance the aesthetic quality of an image by improving its composition through spatial cropping. Previous methods often rely on saliency prediction or retrieval augmentation, ignoring the task's core requirement: a deep understanding of composition and aesthetics. Consequently, saliency-based methods struggle to make compositional trade-offs in complex scenes, while retrieval-based methods blindly refer to similar cases, lacking adaptive reasoning for unique scenes. Both approaches fail to align their automated cropping results with those of human experts. To address the above issues, we propose a novel paradigm that reformulates aesthetic cropping as a multimodal reasoning task, aiming to activate the VLM's analytical and comprehension capabilities in aesthetics. We design a Compositional Reasoning and Optimizing Preference method (CROP) that directs the VLM to think like a professional photographer. It deconstructs a complex and subjective aesthetic problem into an "analysis-proposal-decision" process, reasoning step by step through the analysis of scene elements and compositional principles. Meanwhile, our expert preference alignment module makes the model's decision consistent with human expert aesthetics. Extensive experiments across multiple datasets validate our method's superiority and component effectiveness.

Subjects:	Computer Vision and Pattern Recognition (cs.CV); Artificial Intelligence (cs.AI)
Cite as:	arXiv:2605.12545 [cs.CV]
	(or arXiv:2605.12545v1 [cs.CV] for this version)
	https://doi.org/10.48550/arXiv.2605.12545 arXiv-issued DOI via DataCite (pending registration)

Submission history

From: Zhitong Dong [view email]
[v1] Sat, 9 May 2026 10:21:51 UTC (12,703 KB)

— Originally published at arxiv.org

Continue reading on arxiv.org

CROP: Expert-Aligned Image Cropping via Compositional Reasoning and Optimizing Preference

Quick Take

Key Points

📖 Reader Mode

Submission history

More from arXiv cs.CV

CoReDiT: Spatial Coherence-Guided Token Pruning and Reconstruction for Efficient Diffusion Transformers

ProtoMedAgent: Multimodal Clinical Interpretability via Privacy-Aware Agentic Workflows

Diagnosing and Correcting Concept Omission in Multimodal Diffusion Transformers

Related in this space

Invisible Orchestrators Suppress Protective Behavior and Dissociate Power-Holders: Safety Risks in Multi-Agent LLM Systems

Enhanced and Efficient Reasoning in Large Learning Models

Measuring and Mitigating Toxicity in Large Language Models: A Comprehensive Replication Study