OverLay++: Dense-Overlap Layout-to-Image Generation Dataset
Quick Answer
This paper shows that OverLay++ introduces a new Layout-to-Image dataset with 500K images, averaging 6.6 objects per image, enhancing annotation density by 1.67 times.
Quick Take
This dataset improves state-of-the-art methods, demonstrating the significance of dense, overlap-aware, and caption-rich supervision for controllable image generation.
Key Points
- OverLay++ contains approximately 500K images with complex object interactions.
- The dataset has an average of 6.6 objects per image, exceeding existing datasets.
- Annotations in OverLay++ are over six times longer than current datasets.
- State-of-the-art methods trained on OverLay++ show consistent improvement.
- The dataset generation pipeline produces dense, overlapping object annotations.
Paper Resources
📖 Reader Mode
~2 min readAbstract:Layout-to-Image generation has made substantial progress in spatial and object-level control. However, existing methods still struggle with complex scenes containing many overlapping and interacting objects. We argue that training data is a particular bottleneck: existing datasets lack examples with dense, complex object interactions. To address this gap, we introduce OverLay++, a large-scale Layout-to-Image dataset with structurally complex scenes. OverLay++ contains approximately 500K images with an average of 6.6 objects per image, exceeding existing datasets by 1.67 times in annotation density. Beyond annotation density, OverLay++ provides rich semantic detail with object captions over six times longer than in current datasets. Our dataset generation pipeline is simple and produces dense, overlapping object annotations with rich per-object captions. Across multiple benchmarks, state-of-the-art Layout-to-Image methods trained on the OverLay++ dataset show consistent improvement and faster convergence, demonstrating the importance of dense, overlap-aware, and caption-rich supervision for controllable image generation.
| Comments: | Accepted at NeurIPS 2026, Evaluations & Datasets Track. Project website: this https URL . Dataset: this https URL |
| Subjects: | Computer Vision and Pattern Recognition (cs.CV) |
| Cite as: | arXiv:2610.09071 [cs.CV] |
| (or arXiv:2610.09071v1 [cs.CV] for this version) | |
| https://doi.org/10.48550/arXiv.2610.09071 arXiv-issued DOI via DataCite (pending registration) |
Submission history
From: Divyansh Srivastava [view email]
[v1]
Tue, 6 Oct 2026 20:16:16 UTC (32,290 KB)
— Originally published at arxiv.org
Want this in your inbox every morning?
Daily brief at your local 8am — bilingual EN/中文, free.
More from arXiv cs.CV
See more →ProMoE-FL: Prototype-conditioned Mixture of Experts for Multimodal Federated Learning with Missing Modalities
ProMoE-FL introduces a Prototype-conditioned Mixture-of-Experts framework for multimodal federated learning, effectively addressing missing modalities. It outperforms existing methods on four chest X-ray datasets, demonstrating superior feature synthesis capabilities in both homogeneous and heterogeneous settings.