GHARP: Real-time Gaussian Head Animation from Large-scale Reconstruction Prior
Quick Answer
GHARP introduces a two-stage method for real-time 3D head animation, achieving state-of-the-art quality on the Ava-256 benchmark while running up to 13x faster on an A100 GPU with 8x fewer Gaussians.
Quick Take
The approach separates identity representation and animation, addressing common issues in expression-driven avatars with a body alignment network to enhance clarity and reduce artifacts.
Key Points
- GHARP animates 3D heads from few images and expression signals in real-time.
- The identity stage builds geometry offline, while animation runs on a lightweight network.
- Achieves state-of-the-art results on Ava-256 with significant speed improvements.
- Introduces a body alignment network to resolve ambiguities in body pose.
- Optimized for mobile devices, balancing fidelity and runtime efficiency.
DeepSignal Analysis
What happened
GHARP is a new method for real-time 3D head animation that separates identity representation from animation, achieving high quality and efficiency. It runs up to 13 times faster on an A100 GPU while using 8 times fewer Gaussians, and it addresses common issues in expression-driven avatars with a body alignment network.
Key evidence
- GHARP achieves state-of-the-art quality on the Ava-256 benchmark, indicating its competitive performance in 3D head animation.
- The method operates up to 13 times faster on an A100 GPU compared to previous approaches, demonstrating significant improvements in runtime efficiency.
- It utilizes 8 times fewer Gaussians than earlier methods, which contributes to its lightweight design suitable for mobile devices.
Why it matters
The advancements presented in GHARP could significantly enhance the development of real-time animation technologies, particularly in applications like gaming and virtual reality. By improving both speed and quality, this method may enable more realistic and responsive avatars, which are crucial for immersive user experiences. The integration of a body alignment network also addresses a common limitation in avatar animation, potentially leading to more accurate representations of human expressions and movements.
Paper Resources
📖 Reader Mode
~2 min readAuthors:Ali Benlalah, Sepehr Johari, Patricia Vitoria, Armin Kappeler, Artem Sevastopolsky, Alexander Jung, Gabriele Fanelli, Kevin Mader, Manuel Breitenstein, Claudia Plüss, Jan Rüegg, Simon Biland, Thomas Etterlin, Dmitry Kostiaev, Mathias Deschler, Brian Amberg, Sebastian Martin
Abstract:We present GHARP (Real-time Gaussian Head Animation from Large-scale Reconstruction Prior), a method that animates 3D human heads in real time from a few input images of a subject and a driving expression signal. We decouple the problem into an identity stage that builds a representation of the subject's geometry and appearance offline, and an animation stage that predicts expression-dependent residuals on top of it at runtime. This separation offers a favorable trade-off with respect to fidelity, quality and runtime: the identity stage can be expensive while the animation stage runs a lightweight network, optimized for mobile devices. Our method performs animation in a semantically structured latent space of a pretrained reconstruction model, where expression changes remain spatially contained, making residual prediction efficient. This reconstruction prior provides a consistent spatial layout, allowing fusion of multiple input views into a compact, fixed-size canonical Gaussian representation. While this two-stage design improves the runtime-quality trade-off, it still inherits a problem common to all expression-driven avatar methods: expression codes describe only the face and thus omit body pose and clothing position, making these regions underspecified in the input. The animation network faces an ill-posed mapping and resorts to averaging over conflicting body appearances, producing blur and temporal flicker. We address this with a body alignment network that learns to align the person's body in the target image with the input reference images, removing the ambiguity from the training signal. Our method achieves state-of-the-art quality on the Ava-256 benchmark while running up to 13x faster on an A100 GPU with 8x fewer Gaussians.
| Comments: | Accepted to ACCV 2026. 35 pages (14 main + references + 15 pages supplementary), 14 figures, 17 tables |
| Subjects: | Computer Vision and Pattern Recognition (cs.CV); Machine Learning (cs.LG) |
| Cite as: | arXiv:2610.10945 [cs.CV] |
| (or arXiv:2610.10945v1 [cs.CV] for this version) | |
| https://doi.org/10.48550/arXiv.2610.10945 arXiv-issued DOI via DataCite (pending registration) |
Submission history
From: Armin Kappeler [view email]
[v1]
Wed, 7 Oct 2026 21:50:43 UTC (20,103 KB)
— Originally published at arxiv.org
Want this in your inbox every morning?
Daily brief at your local 8am — bilingual EN/中文, free.
More from arXiv cs.CV
See more →ProMoE-FL: Prototype-conditioned Mixture of Experts for Multimodal Federated Learning with Missing Modalities
ProMoE-FL introduces a Prototype-conditioned Mixture-of-Experts framework for multimodal federated learning, effectively addressing missing modalities. It outperforms existing methods on four chest X-ray datasets, demonstrating superior feature synthesis capabilities in both homogeneous and heterogeneous settings.