Reliability-Aware Monocular Depth Supervision for Sparse-View Neural Reconstruction
Quick Answer
This study explores monocular depth supervision for sparse-view neural reconstruction using Depth Anything V2.
Quick Take
While Mip-NeRF-360 shows marginal gains, Splatfacto improves PSNR from 14.903 to 15.932, highlighting the importance of selective depth supervision in enhancing reconstruction quality.
Key Points
- Sparse-view neural reconstruction faces challenges in outdoor driving scenes.
- Masked monocular depth supervision yields minimal gains for Mip-NeRF-360.
- Splatfacto benefits significantly, improving PSNR and reducing RMSE.
- Reliable low-error regions are crucial for effective depth supervision.
- Depth supervision can degrade RGB rendering quality in strong multi-view scenarios.
Paper Resources
📖 Reader Mode
~2 min readAbstract:Sparse-view neural reconstruction is challenging in outdoor driving scenes, where cameras usually move along a narrow forward-facing trajectory and provide limited multi-view overlap. Although monocular depth estimators can provide dense geometric priors, their predictions are noisy, and not uniformly reliable across image regions. In this work, we study monocular depth supervision for sparse-view neural reconstruction. We use Depth Anything V2 as a dense monocular depth prior, align its predictions to metric depth using scale-shift fitting, and apply depth supervision selectively through photometric masks generated from an RGB-only baseline model. We evaluate this strategy on two representative scene representations: Mip-NeRF-360 and Splatfacto. On KITTISeq02 under an every2 sparse-view setting, masked monocular depth supervision gives only marginal rendering gains for Mip-NeRF-360 and does not improve metric geometry. In contrast, Splatfacto benefits more clearly, improving PSNR from 14.903 to 15.932 and reducing RMSE from 0.542 to 0.100. Additional KITTISeq05 experiments and matched-ratio mask ablations further show that the gains for Splatfacto come from selecting reliable low-error regions rather than simply reducing the number of depth-supervised pixels. Additional experiments on the Bicycle scene show that depth supervision can improve geometry while hurting RGB rendering quality when multi-view coverage is already strong. Overall, our results suggest that monocular depth priors are useful for under-constrained sparse-view reconstruction, but should be applied selectively and with moderate weighting.
| Comments: | 10 pages, 6 figures. All authors contributed equally |
| Subjects: | Computer Vision and Pattern Recognition (cs.CV); Graphics (cs.GR) |
| Cite as: | arXiv:2607.02554 [cs.CV] |
| (or arXiv:2607.02554v1 [cs.CV] for this version) | |
| https://doi.org/10.48550/arXiv.2607.02554 arXiv-issued DOI via DataCite |
Submission history
From: Wei-Teng Chu [view email]
[v1]
Sat, 27 Jun 2026 07:39:22 UTC (3,952 KB)
— Originally published at arxiv.org
Want this in your inbox every morning?
Daily brief at your local 8am — bilingual EN/中文, free.
More from arXiv cs.CV
See more →ProMoE-FL: Prototype-conditioned Mixture of Experts for Multimodal Federated Learning with Missing Modalities
ProMoE-FL introduces a Prototype-conditioned Mixture-of-Experts framework for multimodal federated learning, effectively addressing missing modalities. It outperforms existing methods on four chest X-ray datasets, demonstrating superior feature synthesis capabilities in both homogeneous and heterogeneous settings.