StyleFields: Multi-Scale AdaIN-Modulated Implicit SDFs for Coarse-to-Fine 3D Shape Reconstruction and Editing
Quick Answer
StyleFields introduces a novel DeepSDF architecture for 3D shape reconstruction, enabling controllable geometric style mixing through depth-aware modulation.
Quick Take
This method achieves high-fidelity reconstructions and effective cross-instance hybrids, with applications in automotive aerodynamics for optimizing car designs.
Key Points
- StyleFields uses multi-level Adaptive Instance Normalization for depth-aware modulation.
- Achieves content-style decoupling without part labels or adversarial training.
- Demonstrates practical application in automotive aerodynamics with a drag predictor.
- Delivers consistent gains in ablation studies over injection depth and supervision granularity.
- Enables targeted edits of global form or surface details in 3D models.
Paper Resources
📖 Reader Mode
~2 min readAbstract:We introduce StyleFields, a DeepSDF-based architecture for high-fidelity 3D reconstruction that enables controllable geometric style mixing: the coarse structure of one object can be combined with the fine-scale details of another. The core idea is depth-aware modulation: instead of a single global code, we inject latents via multi-level Adaptive Instance Normalization at several decoder depths, and supervise matching auxiliary heads with a coarse-to-fine schedule while gradually growing network depth. This aligns early layers with global shape and later layers with high-frequency detail, achieving content-style decoupling without part labels or adversarial training. StyleFields delivers faithful reconstructions, convincing cross-instance hybrids, and consistent gains in ablations over injection depth and supervision granularity. We further demonstrate a practical application in automotive aerodynamics: a learned surrogate drag predictor serves as a differentiable objective to optimize reconstructed cars, allowing targeted edits of global form or surface details by freezing the complementary latent stream. StyleFields offers a simple, effective recipe for controllable implicit reconstruction and downstream performance-driven design.
| Comments: | 39 pages, 20 figures, 3 tables. Includes supplementary material |
| Subjects: | Computer Vision and Pattern Recognition (cs.CV); Graphics (cs.GR) |
| Cite as: | arXiv:2610.09200 [cs.CV] |
| (or arXiv:2610.09200v1 [cs.CV] for this version) | |
| https://doi.org/10.48550/arXiv.2610.09200 arXiv-issued DOI via DataCite (pending registration) |
Submission history
From: Ehsan Garaaghaji [view email]
[v1]
Tue, 6 Oct 2026 22:51:51 UTC (17,564 KB)
— Originally published at arxiv.org
Want this in your inbox every morning?
Daily brief at your local 8am — bilingual EN/中文, free.
More from arXiv cs.CV
See more →ProMoE-FL: Prototype-conditioned Mixture of Experts for Multimodal Federated Learning with Missing Modalities
ProMoE-FL introduces a Prototype-conditioned Mixture-of-Experts framework for multimodal federated learning, effectively addressing missing modalities. It outperforms existing methods on four chest X-ray datasets, demonstrating superior feature synthesis capabilities in both homogeneous and heterogeneous settings.