RDGSplat: Render-Dedicated Geometry for Novel View Synthesis
Quick Answer
RDGSplat introduces a framework for novel view synthesis that optimizes rendering geometry from a frozen 3D foundation model, improving WM2.0 scores from 20.918 to 24.266 dB on the RE10K benchmark with 205.5M additional parameters.
Quick Take
This method preserves the original model's metric predictions while enhancing rendering quality across multiple backbones.
Key Points
- RDGSplat decodes dedicated geometry for rendering from a frozen 3D model.
- Improves novel view synthesis across three feed-forward backbones.
- Achieves a WM2.0 score increase from 20.918 to 24.266 dB on RE10K.
- Utilizes 205.5M additional parameters while keeping the backbone weights frozen.
- Maintains original metric predictions during the optimization process.
DeepSignal Analysis
What happened
RDGSplat is a new framework for novel view synthesis that optimizes rendering geometry from a frozen 3D foundation model. It enhances WM2.0 scores on the RE10K benchmark from 20.918 to 24.266 dB while adding 205.5 million parameters. The method maintains the original model's metric predictions and improves rendering quality across various backbones.
Key evidence
- RDGSplat optimizes rendering geometry from a frozen 3D foundation model, improving WM2.0 scores from 20.918 to 24.266 dB on the RE10K benchmark.
- The framework adds 205.5 million parameters while keeping the original model's metric predictions unchanged.
- Extensive experiments demonstrate that RDGSplat enhances novel view synthesis across three feed-forward backbones on four benchmarks.
Why it matters
The development of RDGSplat is significant as it addresses the limitations of existing methods that require updating backbone weights, which can compromise the model's original metric predictions. By preserving these predictions while enhancing rendering quality, RDGSplat offers a more efficient approach to novel view synthesis. This could have implications for various applications in computer vision, particularly in areas requiring high-quality rendering from 3D models.
Paper Resources
📖 Reader Mode
~2 min readAbstract:3D foundation models enable efficient novel view synthesis by carrying a Gaussian head on the representation they already use for reconstruction. However, the views they render fall short of the geometry they recover, because that geometry is estimated under a metric objective and never scored on how it renders. Recent methods alleviate this by updating the backbone weights, but they thereby discard the metric predictions the model was built for and must be repeated for every new backbone. To this end, we propose RDGSplat, a framework that decodes a second geometry dedicated to rendering from a frozen 3D foundation model, leaving its metric predictions intact. In particular, we devise Render-Dedicated Geometry Decoding, which duplicates the pretrained decoders and optimizes the duplicates under photometric supervision alone. Then, a Target-Pose Conditioned Adapter is introduced to reformulate the representation those decoders read, conditioned on the target camera pose rather than the target image. Extensive experiments show that RDGSplat improves novel view synthesis across three feed-forward backbones on four benchmarks, with every pretrained weight frozen. On RE10K, it raises WM2.0 from 20.918 to 24.266\,dB while training 205.5\,M added parameters against a frozen 1.4\,B backbone, and the depth and pose the same model predicts are unchanged.
| Subjects: | Computer Vision and Pattern Recognition (cs.CV) |
| Cite as: | arXiv:2610.09173 [cs.CV] |
| (or arXiv:2610.09173v1 [cs.CV] for this version) | |
| https://doi.org/10.48550/arXiv.2610.09173 arXiv-issued DOI via DataCite (pending registration) |
Submission history
From: Zhijie Zheng [view email]
[v1]
Tue, 6 Oct 2026 22:18:11 UTC (13,803 KB)
— Originally published at arxiv.org
Want this in your inbox every morning?
Daily brief at your local 8am — bilingual EN/中文, free.
More from arXiv cs.CV
See more →ProMoE-FL: Prototype-conditioned Mixture of Experts for Multimodal Federated Learning with Missing Modalities
ProMoE-FL introduces a Prototype-conditioned Mixture-of-Experts framework for multimodal federated learning, effectively addressing missing modalities. It outperforms existing methods on four chest X-ray datasets, demonstrating superior feature synthesis capabilities in both homogeneous and heterogeneous settings.