Depth-to-RGB: Repurposing a Frozen Depth Estimator for Geometry-Guided Compositing
Quick Answer
This paper shows that The Depth-to-RGB (D2R) framework enhances object compositing by predicting composite depth from RGB inputs, achieving a 31.4% reduction in OOD Stage-1 AbsRel and leading benchmarks in geometry and photometric quality.
Quick Take
D2R outperforms 12 open-source and 3 closed-source baselines, reducing AbsRel by 43.7% and improving PSNR by 2.4 dB on category-disjoint data.
Key Points
- D2R predicts composite depth using reference-conditioned corrections to a frozen depth estimator.
- Encoder-feature supervision reduces OOD Stage-1 AbsRel by 31.4% compared to decoded-depth supervision.
- D2R outperforms 12 open-source and 3 closed-source baselines in geometry and photometric quality.
- On category-disjoint data, D2R reduces AbsRel by 43.7% and improves PSNR by 2.4 dB.
- D2R leads in identity metrics and reduces mean CLIP reference cosine distance by 55%.
Paper Resources
📖 Reader Mode
~2 min readAbstract:Reference-based object compositing inserts or replaces an object using a background image, a reference image, and a 2D compositing mask. These inputs guide appearance and placement but leave the completed scene's geometry implicit, which can distort object structure or alter the surroundings. Our Depth-to-RGB (D2R) framework predicts composite depth for a scene not yet observed in the RGB inputs. It learns reference-conditioned corrections to a frozen depth estimator using encoder features of paired completed scenes as targets. The unchanged decoder maps the corrected representation to the intended scene's depth, which a separately trained renderer holds fixed during RGB synthesis. Under matched architecture and training, encoder-feature supervision reduces OOD Stage-1 AbsRel by 31.4% relative to decoded-depth supervision. We also introduce AnyInsertion++ with paired in-distribution and category-disjoint splits to evaluate generalization beyond compositing training categories. The complete D2R system leads 12 open-source and 3 closed-source baselines in estimator-derived geometry and photometric quality on both paired splits. On category-disjoint data, D2R reduces AbsRel by 43.7% and improves PSNR by 2.4 dB over the matched RGB baseline. Across three unpaired benchmarks, D2R leads both identity metrics and reduces mean CLIP reference cosine distance by 55% relative to the strongest baseline. Project page: this https URL
| Subjects: | Computer Vision and Pattern Recognition (cs.CV) |
| Cite as: | arXiv:2610.09125 [cs.CV] |
| (or arXiv:2610.09125v1 [cs.CV] for this version) | |
| https://doi.org/10.48550/arXiv.2610.09125 arXiv-issued DOI via DataCite (pending registration) |
Submission history
From: Sanghyun Jo [view email]
[v1]
Tue, 6 Oct 2026 21:19:26 UTC (33,691 KB)
— Originally published at arxiv.org
Want this in your inbox every morning?
Daily brief at your local 8am — bilingual EN/中文, free.
More from arXiv cs.CV
See more →ProMoE-FL: Prototype-conditioned Mixture of Experts for Multimodal Federated Learning with Missing Modalities
ProMoE-FL introduces a Prototype-conditioned Mixture-of-Experts framework for multimodal federated learning, effectively addressing missing modalities. It outperforms existing methods on four chest X-ray datasets, demonstrating superior feature synthesis capabilities in both homogeneous and heterogeneous settings.