Meta-CRS: A Dynamic Meta-Learning Approach for Effective Conversational Recommender System

被引：3

作者：

Ni, Yuxin ^{[1
]}

Xia, Yunwen ^{[2
]}

Fang, Hui ^{[3
,4
]}

Long, Chong ^{[5
]}

Kong, Xinyu ^{[5
]}

Li, Daqian ^{[5
]}

Yang, Dong ^{[5
]}

Zhang, Jie ^{[2
]}

机构：

[1] Nanyang Technol Univ, 50 Nanyang Ave, Singapore 639798, Singapore

[2] Nanyang Technol Univ, Sch Comp Sci & Engn, 50 Nanyang Ave, Singapore 639798, Singapore

[3] Shanghai Univ Finance & Econ, RIIS, 100 Wudong Rd, Shanghai 200433, Peoples R China

[4] Shanghai Univ Finance & Econ, SIME, 100 Wudong Rd, Shanghai 200433, Peoples R China

[5] Ant Grp, Z Space 556 Xixi Rd, Hangzhou, Peoples R China

来源：

ACM TRANSACTIONS ON INFORMATION SYSTEMS | 2024年 / 42卷 / 01期

基金：

上海市自然科学基金; 中国国家自然科学基金;

关键词：

Conversational recommender system; reinforcement learning; meta learning; prior knowledge; knowledge graph; dynamic graph;

D O I：

10.1145/3604804

中图分类号：

TP [自动化技术、计算机技术];

学科分类号：

0812 ;

摘要：

Conversational recommender system (CRS) enhances the recommender system by acquiring the latest user preference through dialogues, where an agent needs to decide "whether to ask or recommend", "which attributes to ask", and "which items to recommend" in each round. To explore these questions, reinforcement learning is adopted in most CRS frameworks. However, existing studies somewhat ignore to consider the connection between the previous rounds and the current round of the conversation, which might lead to the lack of prior knowledge and inaccurate decisions. In this view, we propose to facilitate the connections between different rounds of conversations in a dialogue session through deep transformer-based multi-channel meta-reinforcement learning, so that the CRS agent can decide each action/decision based on previous states, actions, and their rewards. Besides, to better utilize a user's historical preferences, we propose a more dynamic and personalized graph structure to support the conversation module and the recommendationmodule. Experiment results on five real-world datasets and an online evaluation with real users in an industrial environment validate the improvement of our method over the state-of-the-art approaches and the effectiveness of our designs.

引用

页数：27

共 62 条

[41] Learning to Compare: Relation Network for Few-Shot Learning [J].

Sung, Flood ;

Yang, Yongxin ;

Zhang, Li ;

Xiang, Tao ;

Torr, Philip H. S. ;

Hospedales, Timothy M. .

2018 IEEE/CVF CONFERENCE ON COMPUTER VISION AND PATTERN RECOGNITION (CVPR), 2018, :1199-1208

[42]

Sutton RS, 2018, ADAPT COMPUT MACH LE, P1

[43]

Thrun S, 1998, LEARNING TO LEARN, P3

[44] Dialogue based recommender system that flexibly mixes utterances and recommendations [J].

Tsumita, Daisuke ;

Takagi, Tomohiro .

2019 IEEE/WIC/ACM INTERNATIONAL CONFERENCE ON WEB INTELLIGENCE (WI 2019), 2019, :51-58

[45]

van Hasselt H, 2016, AAAI CONF ARTIF INTE, P2094

[46]

Vaswani A, 2017, ADV NEUR IN, V30

[47]

Velickovic P., 2018, INT C LEARNING REPRE, DOI [DOI 10.48550/ARXIV.1710.10903, 10.48550/arXiv.1710.10903]

[48]

Vinyals Oriol, 2016, Advances in Neural Information Processing Systems, V29

[49]

Wang J. X., 2017, P COGSCI, P1

[50]

Wang Xiaolei, 2022, P 28 ACM SIGKDD C KN

← 1 2 3 4 5 6 7 →