Dynamic Graph Attention for Referring Expression Comprehension

被引:161
作者
Yang, Sibei [1 ]
Li, Guanbin [2 ]
Yu, Yizhou [1 ,3 ]
机构
[1] Univ Hong Kong, Hong Kong, Peoples R China
[2] Sun Yat Sen Univ, Guangzhou, Peoples R China
[3] Deepwise AI Lab, Beijing, Peoples R China
来源
2019 IEEE/CVF INTERNATIONAL CONFERENCE ON COMPUTER VISION (ICCV 2019) | 2019年
基金
中国国家自然科学基金;
关键词
D O I
10.1109/ICCV.2019.00474
中图分类号
TP18 [人工智能理论];
学科分类号
081104 ; 0812 ; 0835 ; 1405 ;
摘要
Referring expression comprehension aims to locate the object instance described by a natural language referring expression in an image. This task is compositional and inherently requires visual reasoning on top of the relationships among the objects in the image. Meanwhile, the visual reasoning process is guided by the linguistic structure of the referring expression. However, existing approaches treat the objects in isolation or only explore the first-order relationships between objects without being aligned with the potential complexity of the expression. Thus it is hard for them to adapt to the grounding of complex referring expressions. In this paper, we explore the problem of referring expression comprehension from the perspective of language-driven visual reasoning, and propose a dynamic graph attention network to perform multi-step reasoning by modeling both the relationships among the objects in the image and the linguistic structure of the expression. In particular, we construct a graph for the image with the nodes and edges corresponding to the objects and their relationships respectively, propose a differential analyzer to predict a language-guided visual reasoning process, and perform stepwise reasoning on top of the graph to update the compound object representation at every node. Experimental results demonstrate that the proposed method can not only significantly surpass all existing state-of-the-art algorithms across three common benchmark datasets, but also generate interpretable visual evidences for stepwisely locating the objects referred to in complex language descriptions.
引用
收藏
页码:4643 / 4652
页数:10
相关论文
共 33 条
[1]   Neural Module Networks [J].
Andreas, Jacob ;
Rohrbach, Marcus ;
Darrell, Trevor ;
Klein, Dan .
2016 IEEE CONFERENCE ON COMPUTER VISION AND PATTERN RECOGNITION (CVPR), 2016, :39-48
[2]  
[Anonymous], 2017, ADV NEURAL INFORM PR
[3]  
[Anonymous], 2018, P IEEE C COMP VIS PA
[4]  
[Anonymous], ICCV
[5]  
Berg T., 2014, EMNLP, P787
[6]  
Burks Arthur W, 1954, MATH TABLES OTHER AI, V8, P53
[7]   Visual Question Reasoning on General Dependency Tree [J].
Cao, Qingxing ;
Liang, Xiaodan ;
Li, Bailin ;
Li, Guanbin ;
Lin, Liang .
2018 IEEE/CVF CONFERENCE ON COMPUTER VISION AND PATTERN RECOGNITION (CVPR), 2018, :7249-7257
[8]  
Deng Chaorui, 2002, P IEEE C COMP VIS PA, P7746
[9]  
Hochreiter S, 1997, Neural Computation, V9, P1735
[10]   Learning to Reason: End-to-End Module Networks for Visual Question Answering [J].
Hu, Ronghang ;
Andreas, Jacob ;
Rohrbach, Marcus ;
Darrell, Trevor ;
Saenko, Kate .
2017 IEEE INTERNATIONAL CONFERENCE ON COMPUTER VISION (ICCV), 2017, :804-813