Devil's on the Edges: Selective Quad Attention for Scene Graph Generation

被引:26
作者
Jung, Deunsol [1 ]
Kim, Sanghyun [1 ]
Kim, Won Hwa [1 ]
Cho, Minsu [1 ]
机构
[1] Pohang Univ Sci & Technol POSTECH, Pohang, Gyeongsangbuk D, South Korea
来源
2023 IEEE/CVF CONFERENCE ON COMPUTER VISION AND PATTERN RECOGNITION (CVPR) | 2023年
关键词
D O I
10.1109/CVPR52729.2023.01790
中图分类号
TP18 [人工智能理论];
学科分类号
081104 ; 0812 ; 0835 ; 1405 ;
摘要
Scene graph generation aims to construct a semantic graph structure from an image such that its nodes and edges respectively represent objects and their relationships. One of the major challenges for the task lies in the presence of distracting objects and relationships in images; contextual reasoning is strongly distracted by irrelevant objects or backgrounds and, more importantly, a vast number of irrelevant candidate relations. To tackle the issue, we propose the Selective Quad Attention Network (SQUAT) that learns to select relevant object pairs and disambiguate them via diverse contextual interactions. SQUAT consists of two main components: edge selection and quad attention. The edge selection module selects relevant object pairs, i.e., edges in the scene graph, which helps contextual reasoning, and the quad attention module then updates the edge features using both edge-to-node and edge-to-edge cross-attentions to capture contextual information between objects and object pairs. Experiments demonstrate the strong performance and robustness of SQUAT, achieving the state of the art on the Visual Genome and Open Images v6 benchmarks.
引用
收藏
页码:18664 / 18674
页数:11
相关论文
共 53 条
[1]   End-to-End Object Detection with Transformers [J].
Carion, Nicolas ;
Massa, Francisco ;
Synnaeve, Gabriel ;
Usunier, Nicolas ;
Kirillov, Alexander ;
Zagoruyko, Sergey .
COMPUTER VISION - ECCV 2020, PT I, 2020, 12346 :213-229
[2]   Knowledge-Embedded Routing Network for Scene Graph Generation [J].
Chen, Tianshui ;
Yu, Weihao ;
Chen, Riquan ;
Lin, Liang .
2019 IEEE/CVF CONFERENCE ON COMPUTER VISION AND PATTERN RECOGNITION (CVPR 2019), 2019, :6156-6164
[3]   Recovering the Unbiased Scene Graphs from the Biased Ones [J].
Chiou, Meng-Jiun ;
Ding, Henghui ;
Yan, Hanshu ;
Wang, Changhu ;
Zimmermann, Roger ;
Feng, Jiashi .
PROCEEDINGS OF THE 29TH ACM INTERNATIONAL CONFERENCE ON MULTIMEDIA, MM 2021, 2021, :1581-1590
[4]  
Choromanski K., 2021, ICLR
[5]   Detecting Visual Relationships with Deep Relational Networks [J].
Dai, Bo ;
Zhang, Yuqi ;
Lin, Dahua .
30TH IEEE CONFERENCE ON COMPUTER VISION AND PATTERN RECOGNITION (CVPR 2017), 2017, :3298-3308
[6]   Learning of Visual Relations: The Devil is in the Tails [J].
Desai, Alakh ;
Wu, Tz-Ying ;
Tripathi, Subarna ;
Vasconcelos, Nuno .
2021 IEEE/CVF INTERNATIONAL CONFERENCE ON COMPUTER VISION (ICCV 2021), 2021, :15384-15393
[7]  
Dosovitskiy A., 2020, ICLR 2021
[8]   Scene Graph Generation with External Knowledge and Image Reconstruction [J].
Gu, Jiuxiang ;
Zhao, Handong ;
Lin, Zhe ;
Li, Sheng ;
Cai, Jianfei ;
Ling, Mingyang .
2019 IEEE/CVF CONFERENCE ON COMPUTER VISION AND PATTERN RECOGNITION (CVPR 2019), 2019, :1969-1978
[9]   From General to Specific: Informative Scene Graph Generation via Balance Adjustment [J].
Guo, Yuyu ;
Gao, Lianli ;
Wang, Xuanhan ;
Hu, Yuxuan ;
Xu, Xing ;
Lu, Xu ;
Shen, Heng Tao ;
Song, Jingkuan .
2021 IEEE/CVF INTERNATIONAL CONFERENCE ON COMPUTER VISION (ICCV 2021), 2021, :16363-16372
[10]   Image Generation from Scene Graphs [J].
Johnson, Justin ;
Gupta, Agrim ;
Li Fei-Fei .
2018 IEEE/CVF CONFERENCE ON COMPUTER VISION AND PATTERN RECOGNITION (CVPR), 2018, :1219-1228