3D FaceShell: Attribute Transfer in 3D Face Avatars as a VLM Defense Mechanism
Quick Answer
This paper shows that The 3D FaceShell framework enhances 3D face avatars' privacy by manipulating VLM interpretations while preserving identity and geometric fidelity.
Quick Take
It uses a learnable Gaussian shell to create subtle perturbations, significantly increasing attribute mismatch rates without compromising visual similarity. Extensive tests on celebrity avatars show its effectiveness against black-box .
Key Points
- 3D FaceShell manipulates VLM interpretations while preserving facial identity.
- Utilizes a learnable Gaussian shell for subtle spatial perturbations.
- Increases attribute injection and mismatch rates significantly.
- Maintains high perceptual similarity in 3D face avatars.
- Demonstrated effectiveness against multiple black-box VLMs.
Paper Resources
📖 Reader Mode
~2 min readAbstract:Photorealistic 3D face avatars are increasingly deployed as reusable digital assets across applications such as telepresence, animation, and personalized media. At the same time, vision-language models (VLMs) can infer sensitive attributes from rendered images with open-ended semantic reasoning without any fine-tuning. This creates a new privacy challenge: once a 3D face avatar is shared, any of its renderings can be analyzed to extract high-level facial attributes. Existing defenses largely operate in 2D image space and do not address identity-preserving semantic manipulation of 3D facial representations. We propose 3D FaceShell, a framework for steering VLM interpretations of faces rendered from 3D models while preserving geometric fidelity and facial identity. 3D FaceShell augments the original 3D representation with a learnable Gaussian shell that produces subtle, spatially distributed perturbations optimized through multi-view embedding alignment. The perturbations are designed to be visually inconspicuous yet sufficient to redirect VLM-based attribute inference in a view-consistent manner. Extensive experiments on reconstructed celebrity face avatars and multiple black-box VLMs demonstrate that 3D FaceShell significantly increases attribute injection and mismatch rates while maintaining high perceptual similarity and identity consistency. Our results show that it is possible to manipulate VLM-level semantic interpretation of 3D faces without compromising their human-recognizable appearance.
| Comments: | ECCV 2026 |
| Subjects: | Computer Vision and Pattern Recognition (cs.CV) |
| Cite as: | arXiv:2607.16280 [cs.CV] |
| (or arXiv:2607.16280v1 [cs.CV] for this version) | |
| https://doi.org/10.48550/arXiv.2607.16280 arXiv-issued DOI via DataCite (pending registration) |
Submission history
From: Weston Bondurant [view email]
[v1]
Thu, 9 Jul 2026 17:55:54 UTC (3,372 KB)
— Originally published at arxiv.org
Want this in your inbox every morning?
Daily brief at your local 8am — bilingual EN/中文, free.
More from arXiv cs.CV
See more →ProMoE-FL: Prototype-conditioned Mixture of Experts for Multimodal Federated Learning with Missing Modalities
ProMoE-FL introduces a Prototype-conditioned Mixture-of-Experts framework for multimodal federated learning, effectively addressing missing modalities. It outperforms existing methods on four chest X-ray datasets, demonstrating superior feature synthesis capabilities in both homogeneous and heterogeneous settings.