Sim-to-Real Transfer of Vision-Language Navigation in Continuous Environments Using an Ackermann-Steered Mobile Robot
Quick Answer
This study presents a novel Vision-Language Navigation (VLN) system that operates in continuous environments using an Ackermann-steered robot, eliminating the need for navigation graphs and panoramic views.
Quick Take
By integrating Cross-Modal Attention and fine-tuning with real-world data, the model demonstrates effective navigation capabilities, achieving robust performance as measured by Success weighted by Path Length (SPL) and Normalized Dynamic Time Warping (nDTW) metrics.
Key Points
- Integrates for natural language-driven navigation.
- Utilizes Cross-Modal Attention architecture trained in simulated environments.
- Fine-tunes with real-world data from a custom-built robot.
- Achieves effective navigation without navigation graphs or panoramic views.
- Evaluated using SPL and nDTW metrics, demonstrating robustness.
Paper Resources
📖 Reader Mode
~2 min readAbstract:Vision-Language Navigation (VLN) enables robots to navigate through environments using natural language instructions, making human-robot interaction intuitive. Traditional VLN models often rely on navigation graphs, 360-degree views, and perfect localization which pose significant challenges when adapting these models to real-world settings. This work addresses these limitations by performing a simulation-to-real domain shift of a VLN approach that operates in continuous environments without requiring navigation graphs or panoramic views. The proposed system integrates vision-language models that align visual inputs and linguistic instructions within a shared embedding space, facilitating natural language-driven navigation. We employ a Cross-Modal Attention (CMA) based architecture trained on an existing dataset in a simulated environment and fine-tune it using real-world data collected from a custom-built Ackermann-steered robot equipped with a camera and a LiDAR sensor. By utilising linear photometric adjustments and fine-tuning on a limited number of episodes, our model successfully adapts to real-world environments, achieving effective navigation while running offline on dedicated hardware. Experimental results, evaluated using Success weighted by Path Length (SPL) and Normalized Dynamic Time Warping (nDTW) metrics, demonstrate the robustness and adaptability of our approach. Keywords: Vision-Language Navigation, Cross-Modal Attention, Natural Language Instructions, Sim-to-Real Transfer, Autonomous Navigation, Ackermann-steering.
| Subjects: | Artificial Intelligence (cs.AI); Robotics (cs.RO) |
| Cite as: | arXiv:2610.07192 [cs.AI] |
| (or arXiv:2610.07192v1 [cs.AI] for this version) | |
| https://doi.org/10.48550/arXiv.2610.07192 arXiv-issued DOI via DataCite (pending registration) |
|
| Related DOI: | https://doi.org/10.1109/ICCAR69571.2026.11549553
DOI(s) linking to related resources |
Submission history
From: Chalindu Abeywansa Nisal [view email]
[v1]
Mon, 5 Oct 2026 18:09:22 UTC (2,321 KB)
— Originally published at arxiv.org
Want this in your inbox every morning?
Daily brief at your local 8am — bilingual EN/中文, free.
More from arXiv cs.AI
See more →HOBA: Hierarchical On-Policy Bidding Agents for Adaptive Online Advertising
HOBA (Hierarchical On-policy Bidding Agents) is a novel hierarchical reinforcement learning framework that enhances online advertising bidding systems by improving adaptability and reducing hyperparameter tuning costs. It utilizes a for hyperparameter inference, a SARSA agent for expert model selection, and a dynamic expert pool for bid execution, achieving a +3.6% increase in target cost during large-scale deployment and outperforming state-of-the-art baselines on AuctionNet.