Closing the Loop: Training-Free Revisit Consistency for Autoregressive Generative Rendering
Quick Answer
The paper presents a training-free method for improving revisit consistency in autoregressive generative rendering, leveraging temporal and spatial correspondences from 3D engine outputs.
Quick Take
It outperforms existing baselines on TartanAir and TartanGround datasets, enhancing video quality without additional training. This approach addresses inconsistencies when the camera revisits locations, crucial for applications in gaming and immersive content.
Key Points
- Introduces a training-free method for autoregressive generative rendering.
- Utilizes temporal and spatial correspondences to enhance revisit consistency.
- Demonstrated on TartanAir and TartanGround datasets.
- Outperforms existing training-free baselines without sacrificing video quality.
- Addresses inconsistencies in long-horizon video generation.
Paper Resources
📖 Reader Mode
~2 min readAbstract:Recent conditional video generation models have shown promising potentials to transform 3D engine renderings, such as depth maps and untextured geometry, into photorealistic videos for gaming and immersive content creation. These applications require long-horizon auto-regressive generation that continuously synthesizes new frames while preserving a persistent 3D world. Auto-regressive generators synthesize video chunk by chunk with a bounded KV cache, so when the camera revisits a location after its context has been evicted, the model often regenerates inconsistent appearance, even though the conditioning renderings (e.g., depth) remain perfectly aligned with the underlying this http URL address this revisit inconsistency without any post-training by exploiting correspondences the 3D engine already provides: temporal correspondence retrieves pose-matched historical latent chunks into the KV cache as loop-closure memory, while spatial correspondence from camera pose and depth reprojection biases token-level attention toward geometrically corresponding regions of the retrieved chunks. We demonstrate our method on loop-closure trajectories mined from TartanAir and TartanGround dataset to mirror complicate real-world application scenarios, where it outperforms existing training-free baselines on revisit consistency without losing overall video quality. Project Page: this https URL
| Subjects: | Computer Vision and Pattern Recognition (cs.CV) |
| Cite as: | arXiv:2607.21848 [cs.CV] |
| (or arXiv:2607.21848v1 [cs.CV] for this version) | |
| https://doi.org/10.48550/arXiv.2607.21848 arXiv-issued DOI via DataCite (pending registration) |
Submission history
From: Wenchao Ma [view email]
[v1]
Thu, 23 Jul 2026 22:33:52 UTC (2,701 KB)
— Originally published at arxiv.org
Want this in your inbox every morning?
Daily brief at your local 8am — bilingual EN/中文, free.
More from arXiv cs.CV
See more →ProMoE-FL: Prototype-conditioned Mixture of Experts for Multimodal Federated Learning with Missing Modalities
ProMoE-FL introduces a Prototype-conditioned Mixture-of-Experts framework for multimodal federated learning, effectively addressing missing modalities. It outperforms existing methods on four chest X-ray datasets, demonstrating superior feature synthesis capabilities in both homogeneous and heterogeneous settings.