HeteroPROPMT: A Real-time and Privacy-Preserving Heterogeneous Collaborative Perception Framework
Quick Answer
HeteroPROMPT is a real-time, privacy-preserving framework for heterogeneous collaborative perception, enhancing autonomous systems' awareness by aligning diverse sensor data with a unified feature space.
Quick Take
It outperforms existing methods in Average Precision on OPV2V-H and V2XSet datasets while using significantly fewer trainable parameters, achieving over 99.99% accuracy in modality classification without exposing proprietary information.
Key Points
- HeteroPROMPT rapidly aligns heterogeneous agent features using modular prompts.
- It maintains frozen encoders while improving collaborative fusion and detection.
- Achieves better Average Precision than state-of-the-art methods with fewer parameters.
- The modality classifier predicts agent modalities with over 99.99% accuracy.
- Designed for metadata-free deployment, enhancing privacy in collaborative perception.
Paper Resources
📖 Reader Mode
~2 min readAbstract:Collaborative Perception (CP) improves autonomous systems' awareness of their surroundings by sharing sensor data, intermediate features, and detection results. In real-world deployments, however, collaborating vehicles often use heterogeneous sensors, perception models, datasets, and training domains, creating feature-space shifts that degrade downstream fusion and detection. Existing approaches typically retrain fusion and detection components or introduce modality-specific feature interpreters. These methods scale poorly to newly joining agents and often require access to proprietary metadata, raising privacy concerns. We propose HeteroPROMPT, a real-time and privacy-preserving framework for heterogeneous collaborative perception. HeteroPROMPT rapidly aligns each heterogeneous agent's features with an ego-centric unified feature space through modular prompts and lightweight learning-based tuning, while keeping agent encoders and the collaborative fusion and detection stacks frozen. Its visual prompt-based training and inference modulate Bird's Eye View (BEV) features across channels and spatial locations with low computational overhead. For metadata-free deployment, an autoencoder learns a compact unified representation and extracts modality cues from shared features, enabling real-time modality classification and routing to the appropriate HeteroPROMPT modules without exposing proprietary agent information. Experiments on the OPV2V-H and V2XSet datasets show that HeteroPROMPT improves Average Precision over state-of-the-art heterogeneous CP methods while using orders of magnitude fewer trainable parameters. This offers a scalable and practical CP solution. The proposed modality classifier also predicts the joining agent's modality from compact features with greater than 99.99 percent accuracy during deployment. Code will be available at this https URL.
| Comments: | Accepted to 2026 IEEE/RSJ International Conference on Intelligent Robots and Systems (IROS). 9 pages, 4 figures, 5 tables |
| Subjects: | Computer Vision and Pattern Recognition (cs.CV); Robotics (cs.RO) |
| Cite as: | arXiv:2607.26283 [cs.CV] |
| (or arXiv:2607.26283v1 [cs.CV] for this version) | |
| https://doi.org/10.48550/arXiv.2607.26283 arXiv-issued DOI via DataCite (pending registration) |
Submission history
From: Armin Maleki [view email]
[v1]
Tue, 28 Jul 2026 21:24:43 UTC (527 KB)
— Originally published at arxiv.org
Want this in your inbox every morning?
Daily brief at your local 8am — bilingual EN/中文, free.
More from arXiv cs.CV
See more →ProMoE-FL: Prototype-conditioned Mixture of Experts for Multimodal Federated Learning with Missing Modalities
ProMoE-FL introduces a Prototype-conditioned Mixture-of-Experts framework for multimodal federated learning, effectively addressing missing modalities. It outperforms existing methods on four chest X-ray datasets, demonstrating superior feature synthesis capabilities in both homogeneous and heterogeneous settings.