SatFix: Absolute Visual Localization of UAVs in Satellite Maps from a Single Oblique Image
Quick Answer
SatFix introduces a novel UAV localization framework that accurately determines position and heading from a single oblique image, achieving 52.08% localization within 50 m and reducing median position error by 34% compared to VGGT-$ ext{Ω}$.
Quick Take
The model operates efficiently on NVIDIA RTX 4090, processing queries in under 0.1 seconds.
Key Points
- SatFix uses satellite-grid features to aggregate UAV visual evidence for localization.
- Achieves median position error of 21.96 m with nine UAV views.
- No explicit 3D map or auxiliary sensors are required for localization.
- Introduces University-Metric for comprehensive evaluation of localization accuracy.
- Median heading error improved from 25.43° to 8.73° with multiple views.
Paper Resources
📖 Reader Mode
~2 min readAbstract:We study absolute metric UAV localization within a provided geo-referenced satellite region, recovering continuous map position and viewing heading from a single oblique image or a short multi-view clip. Existing cross-view geo-localization methods retrieve the most similar satellite tile from a gallery and report Recall@K, but retrieval depends on gallery sampling, provides no heading estimate, and returns a tile index rather than a continuous coordinate. We propose SatFix, a feed-forward UAV--satellite localization framework built on VGGT-$\Omega$. Satellite-grid features act as queries that aggregate UAV visual evidence, and two lightweight heads regress a 3-DoF pose in the satellite-map frame: continuous 2D position and heading. SatFix requires no explicit 3D map, rendered bird's-eye image, auxiliary sensor, or test-time pose alignment. A single model supports both single- and multi-view inputs, with trajectory constraints used during multi-view training. For metric evaluation, we introduce University-Metric, where satellite imagery is re-collected over a region up to 10.7$\times$ longer on a side (about 114$\times$ the ground area) than the original University-1652 tiles, with continuous position and heading labels for the original UAV tours. With one UAV view, SatFix localizes 52.08% of test frames within 50 m and 17.34% within 10 m, with median position and heading errors of 45.66 m and $20.81^\circ$, respectively. Inference takes under 0.1 s per single-view query on an NVIDIA RTX 4090. With nine UAV views, the median position error falls to 21.96 m and the median heading error to $8.73^\circ$. Compared with a fine-tuned VGGT-$\Omega$ baseline, SatFix reduces median position error by 34.0% and nine-view median heading error from $25.43^\circ$ to $8.73^\circ$.
| Subjects: | Computer Vision and Pattern Recognition (cs.CV) |
| Cite as: | arXiv:2610.11049 [cs.CV] |
| (or arXiv:2610.11049v1 [cs.CV] for this version) | |
| https://doi.org/10.48550/arXiv.2610.11049 arXiv-issued DOI via DataCite (pending registration) |
Submission history
From: Jiarui Zeng [view email]
[v1]
Thu, 8 Oct 2026 01:01:46 UTC (1,471 KB)
— Originally published at arxiv.org
Want this in your inbox every morning?
Daily brief at your local 8am — bilingual EN/中文, free.
More from arXiv cs.CV
See more →ProMoE-FL: Prototype-conditioned Mixture of Experts for Multimodal Federated Learning with Missing Modalities
ProMoE-FL introduces a Prototype-conditioned Mixture-of-Experts framework for multimodal federated learning, effectively addressing missing modalities. It outperforms existing methods on four chest X-ray datasets, demonstrating superior feature synthesis capabilities in both homogeneous and heterogeneous settings.