Sim-to-real transfer for Vision-Language Navigation using Cross-Modal Attention on an Ackermann-steered robot (arXiv:2610.07192v1)
This arXiv v1 paper presents a VLN system that operates in continuous environments without navigation graphs or panoramic views, using a Cross-Modal Attention architecture trained in simulation and fine-tuned on limited real-world episodes collected with a custom Ackermann-steered robot outfitted with a camera and LiDAR. The authors apply linear photometric adjustments and demonstrate sim-to-real adaptation with evaluations reported via SPL and nDTW, running offline on dedicated hardware.