Moving Like a Human: Ego-Motion-Normalized Temporal Signatures for Real-Time Aerial Person Tracking on Milliwatt-Class Hardware
Quick Answer
The EMTS-Det system enables real-time aerial person tracking on low-power hardware, achieving 31.85 FPS with 0.462 AP25 on a Raspberry Pi Zero 2W.
Quick Take
It employs ego-motion normalization and a 22k-parameter network, outperforming YOLOv8n in accuracy and speed despite significantly lower computational resources. The system demonstrates high reliability with a 97.9% lock recall and effective occlusion recovery.
Key Points
- EMTS-Det processes frames into ego-motion-normalized channels for improved tracking.
- Achieves 0.694 AP25 in-domain performance on a low-power Raspberry Pi.
- Outperforms YOLOv8n with 1,100 times more compute at 1.95 FPS and 0.172 AP25.
- Utilizes a Kalman filter for stabilized target tracking and occlusion recovery.
- Demonstrates a 97.9% lock recall rate with zero false re-locks during tests.
Paper Resources
📖 Reader Mode
~2 min readAbstract:Follow-me person tracking must run on the drone itself, where affordable companion computers offer only a few effective int8 GFLOP/s. At typical follow distances a person spans 10-60 pixels, indistinguishable from clutter and beyond the reach of single-frame appearance detectors. The missing evidence is temporal and belongs in the input representation, computed analytically, rather than in learned temporal machinery. EMTS-Det is a five-stage system that estimates ego-motion, converts each frame into ego-motion-normalized residual-motion channels, detects person centers with a 22k-parameter, 7.6-MFLOP network, tracks a locked target with a Kalman filter in stabilized coordinates, and verifies tracks with a 1-D convolutional classifier of human motion (ROC AUC 0.941). Training uses a synthetic-motion curriculum with motion channels generated by the deployed ego-motion code. Multi-seed ablations locate the value in generalization: on held-out VisDrone-DET a luminance-only variant collapses to 0.051 AP25 versus 0.415, as does YOLOv8n fine-tuned identically despite 1,100 times the compute, while the deployed int8 detector reaches 0.694 AP25 in-domain and 0.444 on this split. Temporal-shift modules lower accuracy, so the deployed detector is stateless. Silent int8 calibration failures are documented; min-max calibration with propagated caches matches float within 0.008 AP. On a Raspberry Pi Zero 2W the pipeline runs at 31.85 FPS with 0.462 AP25 and 0.714 recall over 1,000 real-world UAV videos, versus 1.95 FPS and 0.172 AP25 for YOLOv8n. A 57-second field sequence shows auto-lock at 1.3 s, 97.9% lock recall, and recovery from all nine occlusions with zero false re-locks.
| Comments: | 20 pages, 10 figures |
| Subjects: | Computer Vision and Pattern Recognition (cs.CV) |
| Cite as: | arXiv:2607.16282 [cs.CV] |
| (or arXiv:2607.16282v1 [cs.CV] for this version) | |
| https://doi.org/10.48550/arXiv.2607.16282 arXiv-issued DOI via DataCite (pending registration) |
Submission history
From: Gholamreza Anbarjafari [view email]
[v1]
Thu, 9 Jul 2026 19:30:00 UTC (1,052 KB)
— Originally published at arxiv.org
Want this in your inbox every morning?
Daily brief at your local 8am — bilingual EN/中文, free.
More from arXiv cs.CV
See more →ProMoE-FL: Prototype-conditioned Mixture of Experts for Multimodal Federated Learning with Missing Modalities
ProMoE-FL introduces a Prototype-conditioned Mixture-of-Experts framework for multimodal federated learning, effectively addressing missing modalities. It outperforms existing methods on four chest X-ray datasets, demonstrating superior feature synthesis capabilities in both homogeneous and heterogeneous settings.