PanoPed: Beyond Bounding Boxes for Sim-to-Real Panoramic Pedestrian Tracking
Quick Answer
PanoPed introduces a sim-to-real benchmark for panoramic pedestrian tracking, featuring 108,000 synthetic frames and 28,002 real frames.
Quick Take
The new Sextant localization head improves the MOTIP baseline from 47.30 to 49.49 HOTA without requiring additional image encoders, demonstrating significant performance gains across all test sequences.
Key Points
- PanoPed-S includes 108,000 synthetic frames from various camera setups.
- PanoPed-R adds 28,002 real frames, with 16,247 densely annotated.
- Sextant, with 0.035M parameters, enhances localization without extra encoders.
- MOTIP baseline improved from 47.30 to 49.49 HOTA using Sextant.
- Synthetic-trained heads also boost performance on real video sequences.
Paper Resources
📖 Reader Mode
~2 min readAbstract:Full-sphere panoramic cameras let fixed monitoring systems and mobile robots track people in every direction, but a planar bounding box does not fully describe where a person is on the sphere. We introduce PanoPed, a sim-to-real benchmark for pedestrian tracking on the full sphere. PanoPed-S contains 108,000 frames from fixed, quadruped-mounted, and drone-mounted cameras, with synchronized masks, depth, camera poses, and 3D pedestrian states. PanoPed-R adds 28,002 real frames from fixed cameras, 16,247 of them densely annotated. We find that an ERP rectangle cannot uniquely determine the spherical center and angular extent of the visible person, while the detector's visual query still carries information about them. Inspired by the sextant's use of angular measurements to locate objects, we propose Sextant, a plug-and-play angular localization head with only about 0.035M parameters. It reuses a frozen detector, keeps track identities unchanged, and needs no extra image encoder. Sextant gives the best result in our PanoPed-S test comparison, raising the strongest baseline, MOTIP, from 47.30 to 49.49 HOTA, with gains on all eight test sequences. Without fine-tuning on real data, the same synthetic-trained heads improve MOTIP and HAT by 0.96-1.14 HOTA on real video, and both seeds improve every real sequence. HAT+Sextant scores best among the compared systems that add no localization image encoder.
| Subjects: | Computer Vision and Pattern Recognition (cs.CV) |
| Cite as: | arXiv:2610.08826 [cs.CV] |
| (or arXiv:2610.08826v1 [cs.CV] for this version) | |
| https://doi.org/10.48550/arXiv.2610.08826 arXiv-issued DOI via DataCite |
Submission history
From: Qinfeng Zhu [view email]
[v1]
Fri, 25 Sep 2026 17:52:19 UTC (7,442 KB)
— Originally published at arxiv.org
Want this in your inbox every morning?
Daily brief at your local 8am — bilingual EN/中文, free.
More from arXiv cs.CV
See more →ProMoE-FL: Prototype-conditioned Mixture of Experts for Multimodal Federated Learning with Missing Modalities
ProMoE-FL introduces a Prototype-conditioned Mixture-of-Experts framework for multimodal federated learning, effectively addressing missing modalities. It outperforms existing methods on four chest X-ray datasets, demonstrating superior feature synthesis capabilities in both homogeneous and heterogeneous settings.