Speaking the Navigator's Language: Trajectory-Grounded Instruction Translation for Frozen Aerial VLN Agents
Quick Answer
This paper shows that The Trajectory-Grounded Instruction Translator (TGIT) improves aerial VLN agents' success rate from 11.33% to 37.93% by translating user intent-driven commands into executable instructions, leveraging trajectory outcomes.
Quick Take
This approach also enhances performance on real human instructions from 11.33% to 32.51% and boosts held-out OpenFly results from 4.95% to 20.79%.
Key Points
- TGIT translates weak inputs into agent-executable commands, improving aerial navigation.
- Success rate for weak inputs increased from 15.27% to 37.93% with TGIT.
- Zero-shot transfer to real human instructions improved success from 11.33% to 32.51%.
- Held-out OpenFly performance rose from 4.95% to 20.79% after TGIT implementation.
- The model addresses the instruction gap in aerial VLN navigation.
Paper Resources
📖 Reader Mode
~2 min readAuthors:Xi Chen, Zhe Liu, Xiaogang Xu, Jiafei Xu, Chunyi Zhou, Yuan Su, Rui Zeng, Tianyu Du, Kelu Yao, Chao Li, Shouling Ji
Abstract:Aerial vision-and-language navigation (VLN) agents are typically trained on detail-rich, trajectory-aligned commands, whereas users issue short, intent-driven instructions; on a frozen OpenFly navigator, this \emph{instruction gap} drops success rate (SR) from $31.03\%$ to $11.33\%$. To scale translator training, we prompt a language model with human-written style examples to convert original commands into paired, intent-centered Weak commands, which yield $15.27\%$ SR. We introduce the \textbf{Trajectory-Grounded Instruction Translator (TGIT)}, a front-end that keeps the navigator frozen and translates Weak inputs into agent-executable commands by learning from its trajectory outcomes. The resulting Weak-trained translator raises Weak-input SR to $37.93\%$ and transfers zero-shot to real human instructions ($11.33\%{\rightarrow}32.51\%$); it also improves held-out OpenFly ($4.95\%{\rightarrow}20.79\%$) and yields recovery on CityNav and AirVLN.
| Subjects: | Artificial Intelligence (cs.AI) |
| Cite as: | arXiv:2610.10635 [cs.AI] |
| (or arXiv:2610.10635v1 [cs.AI] for this version) | |
| https://doi.org/10.48550/arXiv.2610.10635 arXiv-issued DOI via DataCite (pending registration) |
Submission history
From: Xi Chen [view email]
[v1]
Wed, 7 Oct 2026 13:52:37 UTC (316 KB)
— Originally published at arxiv.org
Want this in your inbox every morning?
Daily brief at your local 8am — bilingual EN/中文, free.
More from arXiv cs.AI
See more →HOBA: Hierarchical On-Policy Bidding Agents for Adaptive Online Advertising
HOBA (Hierarchical On-policy Bidding Agents) is a novel hierarchical reinforcement learning framework that enhances online advertising bidding systems by improving adaptability and reducing hyperparameter tuning costs. It utilizes a for hyperparameter inference, a SARSA agent for expert model selection, and a dynamic expert pool for bid execution, achieving a +3.6% increase in target cost during large-scale deployment and outperforming state-of-the-art baselines on AuctionNet.