Depth Estimators Are Implicit Neural Fields for 3D Scene Geometry Inpainting and Reconstruction
Quick Answer
This paper shows that The Neural Depth Field (NDF) model enhances 3D scene geometry inpainting by addressing inconsistencies in depth estimators, achieving a 63.3% reduction in cross-view inconsistency and a 23.1% improvement in inpainting accuracy.
Quick Take
This model adapts to target domains and maintains geometric consistency, outperforming existing methods across diverse scene data.
Key Points
- NDF adapts to target domains using observed depth data.
- Achieves state-of-the-art performance in 3D scene geometry inpainting.
- Reduces cross-view inconsistency by 63.3%.
- Improves inpainting accuracy by 23.1%.
- Demonstrates high fidelity across indoor scans and satellite imagery.
Paper Resources
📖 Reader Mode
~2 min readAbstract:The 3D geometry of real-world scene data is often incomplete. Mainstream methods use depth estimators to inpaint missing structure. However, their prediction results can be inconsistent with observed geometry, or unreliable on out-of-distribution data. To solve these problems, we propose Neural Depth Field (NDF). Our key insight is that a depth estimator can also be a scene-level implicit field. As an estimator, it adapts to the target domain by learning observed depth data. As an implicit field, it fits the existing geometry to maintain consistency. Under this view, NDF addresses both problems through a single test-time optimization. Experiments show that NDF produces high-fidelity and globally consistent geometry across diverse scene data, ranging from indoor scans to satellite imagery. It reduces cross-view inconsistency by 63.3\% and improves inpainting accuracy by 23.1\%, achieving state-of-the-art performance in 3D scene geometry inpainting. The code is available at: this https URL.
| Subjects: | Computer Vision and Pattern Recognition (cs.CV); Artificial Intelligence (cs.AI); Machine Learning (cs.LG) |
| Cite as: | arXiv:2607.16286 [cs.CV] |
| (or arXiv:2607.16286v1 [cs.CV] for this version) | |
| https://doi.org/10.48550/arXiv.2607.16286 arXiv-issued DOI via DataCite (pending registration) |
Submission history
From: Yingzhao Jian [view email]
[v1]
Sat, 11 Jul 2026 12:52:58 UTC (26,690 KB)
— Originally published at arxiv.org
Want this in your inbox every morning?
Daily brief at your local 8am — bilingual EN/中文, free.
More from arXiv cs.CV
See more →ProMoE-FL: Prototype-conditioned Mixture of Experts for Multimodal Federated Learning with Missing Modalities
ProMoE-FL introduces a Prototype-conditioned Mixture-of-Experts framework for multimodal federated learning, effectively addressing missing modalities. It outperforms existing methods on four chest X-ray datasets, demonstrating superior feature synthesis capabilities in both homogeneous and heterogeneous settings.