LeapTalk: Breaking the Latency-Quality Trade-off in Talking Head Generation
Quick Answer
LeapTalk introduces a novel framework for real-time talking-head generation, achieving high-fidelity video output at 200 FPS with a single forward step.
Quick Take
By utilizing a unique data-to-data transport formulation and a heterogeneous distillation framework, it effectively reduces identity drift and enhances temporal stability, outperforming existing methods in efficiency and stability.
Key Points
- Achieves real-time talking-head generation at 200 FPS with a single forward step.
- Introduces a data-to-data transport formulation to mitigate identity drift.
- Utilizes a heterogeneous distillation framework for smooth knowledge transfer.
- Maintains fine-grained lip synchronization under extreme step reduction.
- Demonstrates superior efficiency and stability compared to existing approaches.
Paper Resources
📖 Reader Mode
~2 min readAbstract:Long-form and real-time talking-head generation remains challenging due to a latency-quality trade-off: inefficient multi-step diffusion prohibits streaming generation, whereas real-time autoregressive approaches suffer from error accumulation and identity drift. To address this drawback, we propose LeapTalk, a novel framework that achieves stable and real-time talking-head generation with a single forward step, scaling to arbitrarily long videos. At the heart of our approach lies a single-step bridge distillation scheme. On the one hand, departing from the conventional noise-to-data paradigm, we introduce a data-to-data transport formulation based on a Brownian bridge. Anchored by a persistent reference, this strategy effectively mitigates identity drift and enhances long-term temporal stability. On the other hand, to enable smooth knowledge transfer from a pre-trained diffusion teacher to the student bridge model, we explore a heterogeneous distillation framework with an SNR-aligned time transformation $\Phi(\tau)$, which bridges the functional discrepancy between the two models. Moreover, we propose an audio-driven classifier-free guidance mechanism to maintain fine-grained lip synchronization under extreme step reduction. Extensive experiments demonstrate that our method achieves high-fidelity and temporally consistent video generation with only 1 step at up to 200 FPS, significantly outperforming existing approaches in both efficiency and stability. Project Page: this https URL
| Subjects: | Computer Vision and Pattern Recognition (cs.CV); Sound (cs.SD) |
| Cite as: | arXiv:2608.00079 [cs.CV] |
| (or arXiv:2608.00079v1 [cs.CV] for this version) | |
| https://doi.org/10.48550/arXiv.2608.00079 arXiv-issued DOI via DataCite |
Submission history
From: Rongxiang Zhang [view email]
[v1]
Wed, 29 Jul 2026 14:44:57 UTC (3,045 KB)
— Originally published at arxiv.org
Want this in your inbox every morning?
Daily brief at your local 8am — bilingual EN/中文, free.
More from arXiv cs.CV
See more →ProMoE-FL: Prototype-conditioned Mixture of Experts for Multimodal Federated Learning with Missing Modalities
ProMoE-FL introduces a Prototype-conditioned Mixture-of-Experts framework for multimodal federated learning, effectively addressing missing modalities. It outperforms existing methods on four chest X-ray datasets, demonstrating superior feature synthesis capabilities in both homogeneous and heterogeneous settings.