S2Tok: Streaming 3D Gaussian Reconstruction with Persistent Spatial Tokens
Quick Answer
S2Tok introduces a feed-forward framework for streaming 3D reconstruction, maintaining a persistent scene state with size-adaptive spatial tokens.
Quick Take
It effectively integrates new observations while limiting redundant storage, achieving competitive rendering quality across four benchmarks without caching previous frames.
Key Points
- S2Tok utilizes latent spatial tokens for persistent scene representation.
- A spatially informed transformer integrates incoming observations with existing tokens.
- The learned admission module selectively expands the representation to reduce storage.
- Hierarchical decoder converts the state into non-pixel-aligned 3D Gaussians.
- Experiments show competitive streaming rendering quality across four benchmarks.
Paper Resources
📖 Reader Mode
~2 min readAbstract:Streaming 3D reconstruction requires more than a sequence of geometric predictions: it requires a persistent scene state that can incorporate new evidence and remain renderable as observations arrive. Latent spatial tokens offer a promising representation for this purpose, but constructing them from an image collection leaves open how to maintain them online, where each observation may both revisit known regions and reveal new content. We introduce S2Tok, a feed-forward framework that maintains a size-adaptive, persistent scene state from uncalibrated image streams. Its central idea is to distinguish updates to the existing representation from selective expansion. A spatially informed transformer integrates each incoming observation with the persistent scene tokens, while a learned admission module selectively expands the representation to limit redundant storage. A hierarchical decoder and Gaussian head convert the evolving state into non-pixel-aligned 3D Gaussians, enabling novel-view rendering without caching previous frames. Experiments across four benchmarks demonstrate competitive streaming rendering quality with compact Gaussian representations. These results support latent spatial tokens as a persistent computational state for online 3D reconstruction, combining learned scene updates with explicit Gaussian rendering.
| Comments: | Project Page: this https URL |
| Subjects: | Computer Vision and Pattern Recognition (cs.CV) |
| Cite as: | arXiv:2610.08978 [cs.CV] |
| (or arXiv:2610.08978v1 [cs.CV] for this version) | |
| https://doi.org/10.48550/arXiv.2610.08978 arXiv-issued DOI via DataCite (pending registration) |
Submission history
From: Fang Li [view email]
[v1]
Tue, 6 Oct 2026 18:40:13 UTC (8,486 KB)
— Originally published at arxiv.org
Want this in your inbox every morning?
Daily brief at your local 8am — bilingual EN/中文, free.
More from arXiv cs.CV
See more →ProMoE-FL: Prototype-conditioned Mixture of Experts for Multimodal Federated Learning with Missing Modalities
ProMoE-FL introduces a Prototype-conditioned Mixture-of-Experts framework for multimodal federated learning, effectively addressing missing modalities. It outperforms existing methods on four chest X-ray datasets, demonstrating superior feature synthesis capabilities in both homogeneous and heterogeneous settings.