Sign in the Air to Unlock: An Interface for authentication in Virtual and Augmented Reality Powered by Point-Voxel Cross-Attention Network
Quick Answer
This paper shows that The 'Sign in the Air to Unlock' interface utilizes a point-voxel Cross-Attention Network (PV-Net) for 3D signature authentication in VR/AR, achieving a 2.5% Equal Error Rate on the DeepAirSig dataset and 76% accuracy on ImmAirsig, enhancing user-centric security without disrupting immersion.
Key Points
- PV-Net models local motion dynamics and global spatial structure from 3D trajectories.
- Evaluated on DeepAirSig with 1,800 signatures and ImmAirsig with 880 samples.
- Traditional authentication methods disrupt immersion and require external hardware.
- 3D behavioral interfaces offer seamless, natural interaction for user authentication.
- Potential applications in various immersive technology sectors.
Paper Resources
📖 Reader Mode
~2 min readAbstract:Significant advancement of immersive technologies such as Virtual and Augmented Reality (VR/AR) and their integration into diverse aspects of modern life need authentication interfaces that are secure, intuitive, and compatible with embodied interaction. Traditional methods such as passwords, PINs, and device-based logins, break immersion and rely on external hardware. Recent 3D-specific behavioral approaches, such as hand-gesture, eye-tracking, and electroencephalography (EEG)-based methods, offer promising alternatives but often require specialized sensors or constrain natural movement, limiting usability in dynamic environments. We present Sign in the Air to Unlock, an in-air signature interface that enables users to authenticate by signing naturally in 3D space which is a familiar, personal, and reproducible gesture. To realize this interface, we design a point-voxel Cross-Attention Network (PV-Net) that jointly models local motion dynamics and global spatial structure from 3D trajectories. The model is evaluated on two datasets: the public DeepAirSig dataset (1,800 signatures from 40 users) and ImmAirsig, a new dataset collected using Meta Quest 2 in immersive VR (880 samples from 22 users). PV-Net achieves an Equal Error Rate of 2.5% on DeepAirSig and 76% classification accuracy on ImmAirSig. These findings highlight the potential of 3D behavioral interfaces for seamless, user-centric authentication that merges security with natural interaction in immersive environments.
| Subjects: | Computer Vision and Pattern Recognition (cs.CV); Cryptography and Security (cs.CR); Human-Computer Interaction (cs.HC); Machine Learning (cs.LG) |
| Cite as: | arXiv:2607.01435 [cs.CV] |
| (or arXiv:2607.01435v1 [cs.CV] for this version) | |
| https://doi.org/10.48550/arXiv.2607.01435 arXiv-issued DOI via DataCite (pending registration) |
Submission history
From: Thiru Siddharth [view email]
[v1]
Wed, 1 Jul 2026 19:56:55 UTC (3,743 KB)
— Originally published at arxiv.org
Want this in your inbox every morning?
Daily brief at your local 8am — bilingual EN/中文, free.
More from arXiv cs.CV
See more →ProMoE-FL: Prototype-conditioned Mixture of Experts for Multimodal Federated Learning with Missing Modalities
ProMoE-FL introduces a Prototype-conditioned Mixture-of-Experts framework for multimodal federated learning, effectively addressing missing modalities. It outperforms existing methods on four chest X-ray datasets, demonstrating superior feature synthesis capabilities in both homogeneous and heterogeneous settings.