RepSGG: Novel Representations of Entities and Relationships for Scene Graph Generation

被引:3
|
作者
Liu, Hengyue [1 ]
Bhanu, Bir [1 ]
机构
[1] Univ Calif Riverside, Dept Elect & Comp Engn, Riverside, CA 92521 USA
基金
美国国家科学基金会;
关键词
Feature extraction; Visualization; Semantics; Task analysis; Detectors; Shape; Training; Scene graph generation; visual relationship detection; long-tailed learning; human-Object interaction;
D O I
10.1109/TPAMI.2024.3402143
中图分类号
TP18 [人工智能理论];
学科分类号
081104 ; 0812 ; 0835 ; 1405 ;
摘要
Scene Graph Generation (SGG) has achieved significant progress recently. However, most previous works rely heavily on fixed-size entity representations based on bounding box proposals, anchors, or learnable queries. As each representation's cardinality has different trade-offs between performance and computation overhead, extracting highly representative features efficiently and dynamically is both challenging and crucial for SGG. In this work, a novel architecture called RepSGG is proposed to address the aforementioned challenges, formulating a subject as queries, an object as keys, and their relationship as the maximum attention weight between pairwise queries and keys. With more fine-grained and flexible representation power for entities and relationships, RepSGG learns to sample semantically discriminative and representative points for relationship inference. Moreover, the long-tailed distribution also poses a significant challenge for generalization of SGG. A run-time performance-guided logit adjustment (PGLA) strategy is proposed such that the relationship logits are modified via affine transformations based on run-time performance during training. This strategy encourages a more balanced performance between dominant and rare classes. Experimental results show that RepSGG achieves the state-of-the-art or comparable performance on the Visual Genome and Open Images V6 datasets with fast inference speed, demonstrating the efficacy and efficiency of the proposed methods.
引用
收藏
页码:8018 / 8035
页数:18
相关论文
共 50 条
  • [11] Label Semantic Knowledge Distillation for Unbiased Scene Graph Generation
    Li, Lin
    Xiao, Jun
    Shi, Hanrong
    Wang, Wenxiao
    Shao, Jian
    Liu, An-An
    Yang, Yi
    Chen, Long
    IEEE TRANSACTIONS ON CIRCUITS AND SYSTEMS FOR VIDEO TECHNOLOGY, 2024, 34 (01) : 195 - 206
  • [12] Scene Graph Generation Using Depth, Spatial, and Visual Cues in 2D Images
    Kumar, Aiswarya S.
    Nair, Jyothisha J.
    IEEE ACCESS, 2022, 10 : 1968 - 1978
  • [13] Relationship-Incremental Scene Graph Generation by a Divide-and-Conquer Pipeline With Feature Adapter
    Li, Xuewei
    Zheng, Guangcong
    Yu, Yunlong
    Ji, Naye
    Li, Xi
    IEEE TRANSACTIONS ON IMAGE PROCESSING, 2025, 34 : 678 - 688
  • [14] Neural Belief Propagation for Scene Graph Generation
    Liu, Daqi
    Bober, Miroslaw
    Kittler, Josef
    IEEE TRANSACTIONS ON PATTERN ANALYSIS AND MACHINE INTELLIGENCE, 2023, 45 (08) : 10161 - 10172
  • [15] Explore Contextual Information for 3D Scene Graph Generation
    Liu, Yuanyuan
    Long, Chengjiang
    Zhang, Zhaoxuan
    Liu, Bokai
    Zhang, Qiang
    Yin, Baocai
    Yang, Xin
    IEEE TRANSACTIONS ON VISUALIZATION AND COMPUTER GRAPHICS, 2023, 29 (12) : 5556 - 5568
  • [16] NICEST: Noisy Label Correction and Training for Robust Scene Graph Generation
    Li, Lin
    Xiao, Jun
    Shi, Hanrong
    Zhang, Hanwang
    Yang, Yi
    Liu, Wei
    Chen, Long
    IEEE TRANSACTIONS ON PATTERN ANALYSIS AND MACHINE INTELLIGENCE, 2024, 46 (10) : 6873 - 6888
  • [17] Divide and Conquer: Subset Matching for Scene Graph Generation in Complex Scenes
    Lin, Xin
    Zeng, Jinquan
    Li, Xingquan
    IEEE ACCESS, 2022, 10 : 39069 - 39079
  • [18] Pair Then Relation: Pair-Net for Panoptic Scene Graph Generation
    Wang, Jinghao
    Wen, Zhengyu
    Li, Xiangtai
    Guo, Zujin
    Yang, Jingkang
    Liu, Ziwei
    IEEE TRANSACTIONS ON PATTERN ANALYSIS AND MACHINE INTELLIGENCE, 2024, 46 (12) : 10452 - 10465
  • [19] Spatial–Temporal Knowledge-Embedded Transformer for Video Scene Graph Generation
    Pu, Tao
    Chen, Tianshui
    Wu, Hefeng
    Lu, Yongyi
    Lin, Liang
    IEEE TRANSACTIONS ON IMAGE PROCESSING, 2024, 33 : 556 - 568
  • [20] Multimodal graph inference network for scene graph generation
    Jingwen Duan
    Weidong Min
    Deyu Lin
    Jianfeng Xu
    Xin Xiong
    Applied Intelligence, 2021, 51 : 8768 - 8783