Boosting Efficient Reinforcement Learning for Vision-and-Language Navigation With Open-Sourced LLM

被引：1

作者：

Wang, Jiawei ^{[1
]}

Wang, Teng ^{[2
]}

Cai, Wenzhe ^{[2
]}

Xu, Lele ^{[2
]}

Sun, Changyin ^{[2
,3
]}

机构：

[1] Tongji Univ, Coll Elect & Informat Engn, Shanghai 201804, Peoples R China

[2] Southeast Univ, Sch Automat, Nanjing 210096, Peoples R China

[3] Anhui Univ, Sch Artificial Intelligence, Hefei 230601, Peoples R China

来源：

IEEE ROBOTICS AND AUTOMATION LETTERS | 2025年 / 10卷 / 01期

基金：

中国国家自然科学基金;

关键词：

Navigation; Trajectory; Visualization; Reinforcement learning; Feature extraction; Cognition; Robots; Transformers; Sun; Large language models; Vision-and-language navigation (VLN); large language models; reinforcement learning (RL); attention; discriminator;

D O I：

10.1109/LRA.2024.3511402

中图分类号：

TP24 [机器人技术];

学科分类号：

080202 ; 1405 ;

摘要：

Vision-and-Language Navigation (VLN) requires an agent to navigate in photo-realistic environments based on language instructions. Existing methods typically employ imitation learning to train agents. However, approaches based on recurrent neural networks suffer from poor generalization, while transformer-based methods are too large in scale for practical deployment. In contrast, reinforcement learning (RL) agents can overcome dataset limitations and learn navigation policies that adapt to environment changes. However, without expert trajectories for supervision, agents struggle to learn effective long-term navigation policies from sparse environment rewards. Instruction decomposition enables agents to learn value estimation faster, making agents more efficient in learning VLN tasks. We propose the Decomposing Instructions with Large Language Models for Vision-and-Language Navigation (DILLM-VLN) method, which decomposes complex navigation instructions into simple, interpretable sub-instructions using a lightweight, open-sourced LLM and trains RL agents to complete these sub-instructions sequentially. Based on these interpretable sub-instructions, we introduce the cascaded multi-scale attention (CMA) and a novel multi-modal fusion discriminator (MFD). CMA integrates instruction features at different scales to provide precise textual guidance. MFD combines scene, object, and action information to comprehensively assess the completion of sub-instructions. Experiment results show that DILLM-VLN significantly improves baseline performance, demonstrating its potential for practical applications.

引用

页码：612 / 619

页数：8

共 42 条

[1] Discovering Intrinsic Subgoals for Vision-and-Language Navigation via Hierarchical Reinforcement Learning
Wang, Jiawei
Wang, Teng
Xu, Lele
He, Zichen
Sun, Changyin
IEEE TRANSACTIONS ON NEURAL NETWORKS AND LEARNING SYSTEMS, 2025, 36 (04) : 6516 - 6528
[2] Improved Speaker and Navigator for Vision-and-Language Navigation
Wu, Zongkai
Liu, Zihan
Wang, Ting
Wang, Donglin
IEEE MULTIMEDIA, 2021, 28 (04) : 55 - 63
[3] LLM as Copilot for Coarse-Grained Vision-and-Language Navigation
Qiao, Yanyuan
Liu, Qianyi
Liu, Jiajun
Liu, Jing
Wu, Qi
COMPUTER VISION - ECCV 2024, PT V, 2025, 15063 : 459 - 476
[4] Vision-Language Navigation Policy Learning and Adaptation
Wang, Xin
Huang, Qiuyuan
Celikyilmaz, Asli
Gao, Jianfeng
Shen, Dinghan
Wang, Yuan-Fang
Wang, William Yang
Zhang, Lei
IEEE TRANSACTIONS ON PATTERN ANALYSIS AND MACHINE INTELLIGENCE, 2021, 43 (12) : 4205 - 4216
[5] Visual Perception Generalization for Vision-and-Language Navigation via Meta-Learning
Wang, Ting
Wu, Zongkai
Wang, Donglin
IEEE TRANSACTIONS ON NEURAL NETWORKS AND LEARNING SYSTEMS, 2023, 34 (08) : 5193 - 5199
[6] Learning from Unlabeled 3D Environments for Vision-and-Language Navigation
Chen, Shizhe
Guhur, Pierre-Louis
Tapaswi, Makarand
Schmid, Cordelia
Laptev, Ivan
COMPUTER VISION, ECCV 2022, PT XXXIX, 2022, 13699 : 638 - 655
[7] Reinforced Vision-and-Language Navigation Based on Historical BERT
Zhang, Zixuan
Qi, Shuhan
Zhou, Zihao
Zhang, Jiajia
Yuan, Hao
Wang, Xuan
Wang, Lei
Xiao, Jing
ADVANCES IN SWARM INTELLIGENCE, ICSI 2023, PT II, 2023, 13969 : 427 - 438
[8] HOP plus : History-Enhanced and Order-Aware Pre-Training for Vision-and-Language Navigation
Qiao, Yanyuan
Qi, Yuankai
Hong, Yicong
Yu, Zheng
Wang, Peng
Wu, Qi
IEEE TRANSACTIONS ON PATTERN ANALYSIS AND MACHINE INTELLIGENCE, 2023, 45 (07) : 8524 - 8537
[9] Rule-Based Reinforcement Learning for Efficient Robot Navigation With Space Reduction
Zhu, Yuanyang
Wang, Zhi
Chen, Chunlin
Dong, Daoyi
IEEE-ASME TRANSACTIONS ON MECHATRONICS, 2022, 27 (02) : 846 - 857
[10] Safe-VLN: Collision Avoidance for Vision-and-Language Navigation of Autonomous Robots Operating in Continuous Environments
Yue, Lu
Zhou, Dongliang
Xie, Liang
Zhang, Feitian
Yan, Ye
Yin, Erwei
IEEE ROBOTICS AND AUTOMATION LETTERS, 2024, 9 (06): : 4918 - 4925

← 1 2 3 4 5 →