SPW-Nav: A Streaming Panoramic World Model for Language-Guided Navigation
Quick Answer
SPW-Nav is a novel streaming panoramic world model that generates 2K 360-degree video in real-time, outperforming previous models in camera-following accuracy and video quality.
Quick Take
It interprets movement instructions to create interactive panoramic videos, enabling applications in virtual reality and embodied agent training.
Key Points
- SPW-Nav streams one minute of 2K 360-degree video in real-time.
- It interprets movement instructions as camera motion for interactive experiences.
- Spherical rotation decoupling and pose-aligned conditioning enhance performance.
- SPW-Nav outperforms prior models in camera-following accuracy and video quality.
- The model supports on-the-fly instruction switching for dynamic interactions.
Paper Resources
📖 Reader Mode
~2 min readAbstract:Language-guided panoramic video generation benefits various downstream applications, such as interactive 3D scene exploration, virtual reality experiences, and embodied agent training. Existing panoramic generators follow predefined trajectories, and interactive world models act through low-level actions in perspective views. We propose SPW-Nav, a streaming panoramic world model that understands movement instructions and streams one minute of 2K 360-degree video in real time from a single panorama. SPW-Nav interprets each instruction in the previously generated panorama as camera motion. Spherical rotation decoupling applies rotation exactly on the sphere, pose-aligned conditioning keeps translation inputs bounded over long streams, and a multi-term memory with a few-step generator continues the scene as instructions change. We also build SPW-NavSet, panoramic videos with camera trajectories and verified instructions. Driven by language, SPW-Nav outperforms prior panoramic generators in camera-following accuracy and video quality, and supports on-the-fly instruction switching.
| Comments: | Project page: this https URL |
| Subjects: | Computer Vision and Pattern Recognition (cs.CV) |
| Cite as: | arXiv:2610.08941 [cs.CV] |
| (or arXiv:2610.08941v1 [cs.CV] for this version) | |
| https://doi.org/10.48550/arXiv.2610.08941 arXiv-issued DOI via DataCite (pending registration) |
Submission history
From: Ziqi Cai [view email]
[v1]
Tue, 6 Oct 2026 18:05:45 UTC (36,690 KB)
— Originally published at arxiv.org
Want this in your inbox every morning?
Daily brief at your local 8am — bilingual EN/中文, free.
More from arXiv cs.CV
See more →ProMoE-FL: Prototype-conditioned Mixture of Experts for Multimodal Federated Learning with Missing Modalities
ProMoE-FL introduces a Prototype-conditioned Mixture-of-Experts framework for multimodal federated learning, effectively addressing missing modalities. It outperforms existing methods on four chest X-ray datasets, demonstrating superior feature synthesis capabilities in both homogeneous and heterogeneous settings.