Synthetic and Derived Training Images for Campus Waste Detection: A Multi-Seed Evaluation with YOLOv8n
Quick Answer
The study evaluates the impact of synthetic and derived images on YOLOv8n's performance for campus waste detection, finding no configuration surpassed the real-only model's mAP@0.5 of 0.691.
Quick Take
Various augmentation strategies yielded mixed results, with background replacement reducing performance significantly. The small test set limits strong class-specific conclusions.
Key Points
- Real dataset: 148 campus photos with 86 for training, 31 for validation, and 31 for testing.
- Synthetic images did not improve YOLOv8n beyond the real-only model's mAP@0.5 of 0.691.
- Background replacement led to a significant drop in performance, reducing mean mAP to 0.560.
- Hand-composite experiments showed no reliable effect on detection performance.
- Small test set size limits the ability to draw strong class-specific conclusions.
Paper Resources
📖 Reader Mode
~2 min readAbstract:Incorrect disposal can contaminate campus recycling streams, and a bin-mounted camera could provide feedback as an item is discarded. We evaluated whether synthetic and derived images improve a YOLOv8n detector for this view. The real dataset contained 148 campus photographs: 86 for training, 31 for validation, and 31 for testing. Twelve joint-training configurations varied the amount and source of added images. We repeated seven principal settings with four matched seeds and computed bootstrap percentile intervals over those seeds. The real-only model reached a mean mAP@0.5 of 0.691 [0.665, 0.722]. Background replacement reduced the mean to 0.560 [0.499, 0.619], isolated-object images gave 0.680 [0.644, 0.724], and the full augmentation pool gave 0.487 [0.438, 0.537]. We also tested hand-and-forearm composites because every real photo showed a held object. Two cutouts in the initial composite set came from test photographs, so we discarded that experiment, rebuilt the set with training-split cutouts, and reran all four seeds. The corrected paired difference was +0.034 [-0.063, 0.199], which does not support a reliable hand-composite effect. Single-seed transfer experiments produced source-dependent rankings between joint mixing and sequential pretraining. None of the evaluated configurations exceeded the real-only baseline. The reported intervals quantify seed variation; the 31-photo test set remains too small for strong class-specific conclusions.
| Subjects: | Computer Vision and Pattern Recognition (cs.CV) |
| ACM classes: | I.4.8; I.2.10 |
| Cite as: | arXiv:2607.19535 [cs.CV] |
| (or arXiv:2607.19535v1 [cs.CV] for this version) | |
| https://doi.org/10.48550/arXiv.2607.19535 arXiv-issued DOI via DataCite (pending registration) |
Submission history
From: Ali Behbahani [view email]
[v1]
Tue, 21 Jul 2026 19:30:57 UTC (5,751 KB)
— Originally published at arxiv.org
Want this in your inbox every morning?
Daily brief at your local 8am — bilingual EN/中文, free.
More from arXiv cs.CV
See more →ProMoE-FL: Prototype-conditioned Mixture of Experts for Multimodal Federated Learning with Missing Modalities
ProMoE-FL introduces a Prototype-conditioned Mixture-of-Experts framework for multimodal federated learning, effectively addressing missing modalities. It outperforms existing methods on four chest X-ray datasets, demonstrating superior feature synthesis capabilities in both homogeneous and heterogeneous settings.