Personalize at Test Time: Learning User Preferences for Image Generation
Quick Answer
This study presents a novel approach for personalizing image generation using diffusion models by learning user preferences from historical image pairs.
Quick Take
The method achieves 77% accuracy in predicting pairwise preferences while simplifying user adaptation through low-dimensional weight optimization, enabling efficient personalization with limited feedback.
Key Points
- Utilizes an autoencoder to compress visual attributes into 50 preference dimensions.
- Employs the Bradley-Terry model for estimating user-specific preference weights.
- Achieves 77% accuracy in predicting pairwise user preferences.
- Simplifies personalization by optimizing a low-dimensional weight vector.
- Guides image generation at inference time while keeping the diffusion model static.
Paper Resources
📖 Reader Mode
~2 min readAbstract:Diffusion models can generate high-quality images, yet aligning their outputs with individual user preferences remains challenging. A key bottleneck is accurately modeling diverse user preferences from limited feedback. Existing approaches often rely on labor-intensive manual preference annotations or vision-language models (VLM) to extract preference information from user interaction histories, introducing substantial annotation or computational costs that limit scalability. We propose an approach that learns personalized reward models directly from users' historical image preference pairs. First, we use an autoencoder to compress hundreds of visual attributes into 50 attribute-anchored preference dimensions and train an evaluator to score images along these dimensions. We then represent each user's preferences as a linear combination of the shared dimension scores, estimating the user-specific weights by maximizing the likelihood of their observed pairwise preferences under the Bradley-Terry model. This formulation reduces per-user adaptation to optimizing a low-dimensional weight vector, simplifying optimization and enabling data-efficient personalization from sparse feedback. The learned personalized rewards guide image generation at inference time while keeping the diffusion model frozen. Experiments on real-user preference data show that our approach achieves approximately 77% held-out pairwise preference prediction accuracy and improves the alignment of generated images with individual user preferences.
| Subjects: | Computer Vision and Pattern Recognition (cs.CV); Artificial Intelligence (cs.AI) |
| Cite as: | arXiv:2610.09015 [cs.CV] |
| (or arXiv:2610.09015v1 [cs.CV] for this version) | |
| https://doi.org/10.48550/arXiv.2610.09015 arXiv-issued DOI via DataCite (pending registration) |
Submission history
From: Jiamu Bai [view email]
[v1]
Tue, 6 Oct 2026 19:10:08 UTC (25,987 KB)
— Originally published at arxiv.org
Want this in your inbox every morning?
Daily brief at your local 8am — bilingual EN/中文, free.
More from arXiv cs.CV
See more →ProMoE-FL: Prototype-conditioned Mixture of Experts for Multimodal Federated Learning with Missing Modalities
ProMoE-FL introduces a Prototype-conditioned Mixture-of-Experts framework for multimodal federated learning, effectively addressing missing modalities. It outperforms existing methods on four chest X-ray datasets, demonstrating superior feature synthesis capabilities in both homogeneous and heterogeneous settings.