CAPE-T2V: Captioner-Anchored Prompt Enhancement toward Two-Sided Conditioning Alignment in Text-to-Video Generation
Quick Answer
CAPE-T2V introduces a two-step framework for enhancing prompt alignment in text-to-video generation, significantly reducing the PE-Caption gap.
Quick Take
It fine-tunes the prompt enhancer and diffusion transformers, achieving superior performance on benchmarks like StoryEval and VBench-2.0, with closer distributions in captions during inference. This approach demonstrates a more effective alignment between training and inference outputs.
Key Points
- CAPE-T2V constructs three types of training examples for prompt enhancement.
- Fine-tuning improves alignment between user prompts and video captions.
- Achieves higher scores on benchmarks like StoryEval and T2V-CompBench.
- Demonstrates a smaller PE-Caption gap compared to baseline methods.
- Project available for further exploration and implementation.
DeepSignal Analysis
What happened
CAPE-T2V is a new framework designed to improve prompt alignment in text-to-video generation. It reduces the PE-Caption gap by fine-tuning both the prompt enhancer and diffusion transformers, leading to better performance on benchmarks like StoryEval and VBench-2.0.
Key evidence
- CAPE-T2V introduces a two-step framework that constructs three types of training examples for the prompt enhancer, improving alignment with user prompts.
- The framework fine-tunes the diffusion transformers on video-derived captions rewritten by the prompt enhancer, achieving higher scores on benchmarks such as StoryEval and VBench-2.0.
- CAPE-T2V demonstrates a smaller PE-Caption gap compared to a baseline, with its fine-tuned captions showing closer distribution to inference-time outputs.
Why it matters
The development of CAPE-T2V addresses a significant challenge in text-to-video generation, where discrepancies between training and inference outputs can hinder performance. By effectively aligning prompts and captions, this framework could enhance the quality of generated videos, making them more relevant and coherent. This advancement may have implications for various applications in media and entertainment, where accurate video generation from text is increasingly sought after.
Paper Resources
📖 Reader Mode
~2 min readAbstract:Text-to-video (T2V) diffusion transformers (DiTs) are trained with detailed video captions, whereas inference often relies on user prompts rewritten by a prompt enhancer (PE). Prior work has improved generation by optimizing the PE, the DiT, or both; some methods have also sought to narrow the training-inference mismatch through shared schemas. Yet even within a shared schema, inference-time PE outputs and DiT training captions may still differ in detail selection, information organization, descriptive granularity, and phrasing. We refer to this residual mismatch as the PE-Caption gap and introduce CAPE-T2V, a two-step Captioner-Anchored Prompt Enhancement framework toward two-sided conditioning alignment in T2V generation. First, CAPE-T2V constructs three types of PE training examples, pairing captioner-generated targets with concise source captions, detailed source captions, or pseudo user prompts derived from those targets. It then fine-tunes the PE to map each input to its paired target. Second, CAPE-T2V fine-tunes the DiT on video-derived captions rewritten by the Anchored PE; the same PE rewrites user prompts at inference. Relative to a baseline using the same caption schema, CAPE-T2V achieves higher aggregate scores on StoryEval, VBench-2.0, and T2V-CompBench across Wan2.2 and LTX-2.3. Further, CAPE-T2V exhibits a smaller PE-Caption gap than the baseline: its DiT fine-tuning captions are closer in distribution to inference-time PE outputs, as measured by squared maximum mean discrepancy in a fixed embedding space. Overall, these results support CAPE-T2V as an effective approach to mitigating the PE-Caption gap. The project is available at this https URL.
| Comments: | Includes appendix; 11 figures. Project page: this https URL |
| Subjects: | Computer Vision and Pattern Recognition (cs.CV) |
| Cite as: | arXiv:2608.03046 [cs.CV] |
| (or arXiv:2608.03046v1 [cs.CV] for this version) | |
| https://doi.org/10.48550/arXiv.2608.03046 arXiv-issued DOI via DataCite (pending registration) |
Submission history
From: Yizhuo Jia [view email]
[v1]
Tue, 4 Aug 2026 02:55:10 UTC (20,004 KB)
— Originally published at arxiv.org
Want this in your inbox every morning?
Daily brief at your local 8am — bilingual EN/中文, free.
More from arXiv cs.CV
See more →ProMoE-FL: Prototype-conditioned Mixture of Experts for Multimodal Federated Learning with Missing Modalities
ProMoE-FL introduces a Prototype-conditioned Mixture-of-Experts framework for multimodal federated learning, effectively addressing missing modalities. It outperforms existing methods on four chest X-ray datasets, demonstrating superior feature synthesis capabilities in both homogeneous and heterogeneous settings.