Pair Then Relation: Pair-Net for Panoptic Scene Graph Generation

被引：6

作者：

Wang, Jinghao ^{[1
]}

Wen, Zhengyu ^{[1
]}

Li, Xiangtai ^{[1
]}

Guo, Zujin ^{[1
]}

Yang, Jingkang ^{[1
]}

Liu, Ziwei ^{[1
]}

机构：

[1] Nanyang Technol Univ, S Lab, Singapore 308232, Singapore

来源：

IEEE TRANSACTIONS ON PATTERN ANALYSIS AND MACHINE INTELLIGENCE | 2024年 / 46卷 / 12期

关键词：

Task analysis; Annotations; Pipelines; Decoding; Visualization; Training; Sparse matrices; Detection transformer; panoptic segmentation; scene graph generation;

D O I：

10.1109/TPAMI.2024.3442301

中图分类号：

TP18 [人工智能理论];

学科分类号：

081104 ; 0812 ; 0835 ; 1405 ;

摘要：

Panoptic Scene Graph (PSG) is a challenging task in Scene Graph Generation (SGG) that aims to create a more comprehensive scene graph representation using panoptic segmentation instead of boxes. Compared to SGG, PSG has several challenging problems: pixel-level segment outputs and full relationship exploration (It also considers thing and stuff relation). Thus, current PSG methods have limited performance, which hinders downstream tasks or applications. This work aims to design a novel and strong baseline for PSG. To achieve that, we first conduct an in-depth analysis to identify the bottleneck of the current PSG models, finding that inter-object pair-wise recall is a crucial factor that was ignored by previous PSG methods. Based on this and the recent query-based frameworks, we present a novel framework: Pair then Relation (Pair-Net), which uses a Pair Proposal Network (PPN) to learn and filter sparse pair-wise relationships between subjects and objects. Moreover, we also observed the sparse nature of object pairs for both. Motivated by this, we design a lightweight Matrix Learner within the PPN, which directly learns pair-wised relationships for pair proposal generation. Through extensive ablation and analysis, our approach significantly improves upon leveraging the segmenter solid baseline. Notably, our method achieves over 10% absolute gains compared to our baseline, PSGFormer.

引用

页码：10452 / 10465

页数：14

共 79 条

[1] Exploring Long Tail Visual Relationship Recognition with Large Vocabulary [J].

Abdelkarim, Sherif ;

Agarwal, Aniket ;

Achlioptas, Panos ;

Chen, Jun ;

Huang, Jiaji ;

Li, Boyang ;

Church, Kenneth ;

Elhoseiny, Mohamed .

2021 IEEE/CVF INTERNATIONAL CONFERENCE ON COMPUTER VISION (ICCV 2021), 2021, :15901-15910

[2] Image Understanding using vision and reasoning through Scene Description Graph [J].

Aditya, Somak ;

Yang, Yezhou ;

Baral, Chitta ;

Aloimonos, Yiannis ;

Fermueller, Cornelia .

COMPUTER VISION AND IMAGE UNDERSTANDING, 2018, 173 :33-45

[3] End-to-End Object Detection with Transformers [J].

Carion, Nicolas ;

Massa, Francisco ;

Synnaeve, Gabriel ;

Usunier, Nicolas ;

Kirillov, Alexander ;

Zagoruyko, Sergey .

COMPUTER VISION - ECCV 2020, PT I, 2020, 12346 :213-229

[4] A Comprehensive Survey of Scene Graphs: Generation and Application [J].

Chang, Xiaojun ;

Ren, Pengzhen ;

Xu, Pengfei ;

Li, Zhihui ;

Chen, Xiaojiang ;

Hauptmann, Alex .

IEEE TRANSACTIONS ON PATTERN ANALYSIS AND MACHINE INTELLIGENCE, 2023, 45 (01) :1-26

[5] Reformulating HOI Detection as Adaptive Set Prediction [J].

Chen, Mingfei ;

Liao, Yue ;

Liu, Si ;

Chen, Zhiyuan ;

Wang, Fei ;

Qian, Chen .

2021 IEEE/CVF CONFERENCE ON COMPUTER VISION AND PATTERN RECOGNITION, CVPR 2021, 2021, :9000-9009

[6] Say As You Wish: Fine-grained Control of Image Caption Generation with Abstract Scene Graphs [J].

Chen, Shizhe ;

Jin, Qin ;

Wang, Peng ;

Wu, Qi .

2020 IEEE/CVF CONFERENCE ON COMPUTER VISION AND PATTERN RECOGNITION (CVPR 2020), 2020, :9959-9968

[7] Knowledge-Embedded Routing Network for Scene Graph Generation [J].

Chen, Tianshui ;

Yu, Weihao ;

Chen, Riquan ;

Lin, Liang .

2019 IEEE/CVF CONFERENCE ON COMPUTER VISION AND PATTERN RECOGNITION (CVPR 2019), 2019, :6156-6164

[8]

Cheng B, 2021, ADV NEUR IN, V34

[9] Masked-attention Mask Transformer for Universal Image Segmentation [J].

Cheng, Bowen ;

Misra, Ishan ;

Schwing, Alexander G. ;

Kirillov, Alexander ;

Girdhar, Rohit .

2022 IEEE/CVF CONFERENCE ON COMPUTER VISION AND PATTERN RECOGNITION (CVPR 2022), 2022, :1280-1289

[10] Panoptic-DeepLab: A Simple, Strong, and Fast Baseline for Bottom-Up Panoptic Segmentation [J].

Cheng, Bowen ;

Collins, Maxwell D. ;

Zhu, Yukun ;

Liu, Ting ;

Huang, Thomas S. ;

Adam, Hartwig ;

Chen, Liang-Chieh .

2020 IEEE/CVF CONFERENCE ON COMPUTER VISION AND PATTERN RECOGNITION (CVPR 2020), 2020, :12472-12482

← 1 2 3 4 5 6 7 8 →