Data-efficient model-based reinforcement learning with trajectory discrimination

被引：0

作者：

Qu, Tuo ^{[1
]}

Duan, Fuqing ^{[1
]}

Zhang, Junge ^{[2
]}

Zhao, Bo ^{[3
]}

Huang, Wenzhen ^{[2
]}

机构：

[1] Beijing Normal Univ, Sch Artificial Intelligence, 19 Xinjiekou Outer St, Beijing 100875, Peoples R China

[2] Chinese Acad Sci, Inst Automat, 95 Zhongguancun East Rd, Beijing 100190, Peoples R China

[3] Beijing Normal Univ, Sch Syst Sci, 19 Xinjiekou Outer St, Beijing 100875, Peoples R China

来源：

COMPLEX & INTELLIGENT SYSTEMS | 2024年 / 10卷 / 02期

关键词：

Reinforcement learning; Deep learning; Continuous control task; World model; OBJECTIVE PENALTY-FUNCTION; PREDICTIVE CONTROL; TRACKING; OPTIMIZATION;

D O I：

10.1007/s40747-023-01247-5

中图分类号：

TP18 [人工智能理论];

学科分类号：

081104 ; 0812 ; 0835 ; 1405 ;

摘要：

Deep reinforcement learning has always been used to solve high-dimensional complex sequential decision problems. However, one of the biggest challenges for reinforcement learning is sample efficiency, especially for high-dimensional complex problems. Model-based reinforcement learning can solve the problem with a learned world model, but the performance is limited by the imperfect world model, so it usually has worse approximate performance than model-free reinforcement learning. In this paper, we propose a novel model-based reinforcement learning algorithm called World Model with Trajectory Discrimination (WMTD). We learn the representation of temporal dynamics information by adding a trajectory discriminator to the world model, and then compute the weight of state value estimation based on the trajectory discriminator to optimize the policy. Specifically, we augment the trajectories to generate negative samples and train a trajectory discriminator that shares the feature extractor with the world model. Experimental results demonstrate that our method improves the sample efficiency and achieves state-of-the-art performance on DeepMind control tasks.

引用

页码：1927 / 1936

页数：10

共 50 条

[1] Data-efficient model-based reinforcement learning with trajectory discrimination
Tuo Qu
Fuqing Duan
Junge Zhang
Bo Zhao
Wenzhen Huang
Complex & Intelligent Systems, 2024, 10 : 1927 - 1936
[2] DATA-EFFICIENT MODEL-BASED REINFORCEMENT LEARNING FOR ROBOT CONTROL
Sun, Ming
Gao, Yue
Liu, Wei
Li, Shaoyuan
INTERNATIONAL JOURNAL OF ROBOTICS & AUTOMATION, 2021, 36 (04): : 211 - 218
[3] Identifying Ordinary Differential Equations for Data-efficient Model-based Reinforcement Learning
Nagel, Tobias
Huber, Marco F.
2024 INTERNATIONAL JOINT CONFERENCE ON NEURAL NETWORKS, IJCNN 2024, 2024,
[4] A Safe and Data-Efficient Model-Based Reinforcement Learning System for HVAC Control
Ding, Xianzhong
An, Zhiyu
Rathee, Arya
Du, Wan
IEEE INTERNET OF THINGS JOURNAL, 2025, 12 (07): : 8014 - 8032
[5] Data-Efficient Task Generalization via Probabilistic Model-Based Meta Reinforcement Learning
Bhardwaj, Arjun
Rothfuss, Jonas
Sukhija, Bhavya
As, Yarden
Hutter, Marco
Coros, Stelian
Krause, Andreas
IEEE ROBOTICS AND AUTOMATION LETTERS, 2024, 9 (04) : 3918 - 3925
[6] Data-efficient Deep Reinforcement Learning for Vehicle Trajectory Control
Frauenknecht, Bernd
Ehlgen, Tobias
Trimpe, Sebastian
2023 IEEE 26TH INTERNATIONAL CONFERENCE ON INTELLIGENT TRANSPORTATION SYSTEMS, ITSC, 2023, : 894 - 901
[7] Model-Based Reinforcement Learning With Probabilistic Ensemble Terminal Critics for Data-Efficient Control Applications
Park, Jonghyeok
Jeon, Soo
Han, Soohee
IEEE TRANSACTIONS ON INDUSTRIAL ELECTRONICS, 2024, 71 (08) : 9470 - 9479
[8] Model-Based Data-Efficient Reinforcement Learning for Active Pantograph Control in High-Speed Railways
Wang, Hui
Liu, Zhigang
Wang, Xufan
Meng, Xiangyu
Wu, Yanbo
Han, Zhiwei
IEEE TRANSACTIONS ON TRANSPORTATION ELECTRIFICATION, 2024, 10 (02): : 2701 - 2712
[9] Data-Efficient Hierarchical Reinforcement Learning
Nachum, Ofir
Gu, Shixiang
Lee, Honglak
Levine, Sergey
ADVANCES IN NEURAL INFORMATION PROCESSING SYSTEMS 31 (NIPS 2018), 2018, 31
[10] Data-Efficient Reinforcement Learning with Probabilistic Model Predictive Control
Kamthe, Sanket
Deisenroth, Marc Peter
INTERNATIONAL CONFERENCE ON ARTIFICIAL INTELLIGENCE AND STATISTICS, VOL 84, 2018, 84

← 1 2 3 4 5 →