ARMS: Anchor-Relational Motion Streaming for Seamless Solo-Social Motion Transitions
Quick Answer
ARMS introduces a novel Anchor-Relational Motion Streaming framework that enhances human motion generation by seamlessly integrating solo and social interactions.
Quick Take
This model outperforms traditional methods in transition smoothness and social coherence, achieving competitive results on human-human interaction benchmarks.
Key Points
- ARMS decouples individual motion evolution from inter-person alignment for better stability.
- Utilizes a causal relational diffusion model to refine motion based on past context.
- Mode-aware relational gating allows for flexible activation of cross-agent connections.
- Demonstrates improved transition smoothness compared to interaction-centric baselines.
- Achieves competitive performance on human-human interaction benchmarks.
Paper Resources
📖 Reader Mode
~2 min readAbstract:Generating temporally continuous and socially coherent human motion from text remains a fundamental challenge, particularly in realistic streams where people act alone, enter interactions, and later disengage. Most existing methods generate fixed-length motion clips under static agent configurations, which makes them brittle to solo-social transitions and unsuitable for incremental generation over long horizons. We propose ARMS, an Anchor-Relational Motion Streaming framework that unifies solo motion and human-human interaction within a single causal generative process. ARMS introduces a dynamics-asymmetric representation that decouples per-person temporal evolution from inter-person alignment via a partner-referenced relative-translation term, enabling seamless switching of social coupling without sacrificing long-horizon stability or spatial consistency between agents. On top of a causal latent space, a causal relational diffusion model progressively refines motion segment by segment using only past context, capturing both intra-person temporal dependencies and inter-person relations. Mode-aware relational gating activates or masks cross-agent connections, allowing the same model to support both solo and interaction generation. Experiments show that ARMS improves transition smoothness and social coherence compared to interaction-centric baselines, while also achieving competitive results on human-human interaction benchmarks.
| Comments: | Accepted by ECCV 2026. Project page: this https URL |
| Subjects: | Computer Vision and Pattern Recognition (cs.CV) |
| Cite as: | arXiv:2607.05733 [cs.CV] |
| (or arXiv:2607.05733v1 [cs.CV] for this version) | |
| https://doi.org/10.48550/arXiv.2607.05733 arXiv-issued DOI via DataCite (pending registration) |
Submission history
From: Huakun Liu [view email]
[v1]
Tue, 7 Jul 2026 01:37:38 UTC (14,278 KB)
— Originally published at arxiv.org
Want this in your inbox every morning?
Daily brief at your local 8am — bilingual EN/中文, free.
More from arXiv cs.CV
See more →ProMoE-FL: Prototype-conditioned Mixture of Experts for Multimodal Federated Learning with Missing Modalities
ProMoE-FL introduces a Prototype-conditioned Mixture-of-Experts framework for multimodal federated learning, effectively addressing missing modalities. It outperforms existing methods on four chest X-ray datasets, demonstrating superior feature synthesis capabilities in both homogeneous and heterogeneous settings.