NICEST: Noisy Label Correction and Training for Robust Scene Graph Generation

被引:8
作者
Li, Lin [1 ]
Xiao, Jun [1 ]
Shi, Hanrong [1 ]
Zhang, Hanwang [2 ]
Yang, Yi [1 ]
Liu, Wei [3 ]
Chen, Long [4 ]
机构
[1] Zhejiang Univ, Coll Comp Sci, Hangzhou 310027, Peoples R China
[2] Nanyang Technol Univ, Sch Comp Sci & Engn, Singapore 639798, Singapore
[3] Tencent, Shenzhen 518000, Peoples R China
[4] Hong Kong Univ Sci & Technol, Dept Comp Sci & Engn, Clear Water Bay, Hong Kong, Peoples R China
基金
中国国家自然科学基金;
关键词
Noise measurement; Training; Visualization; Annotations; NIST; Benchmark testing; Task analysis; Multi-Teacher knowledge distillation; noisy label learning; out-of-distribution; scene graph generation;
D O I
10.1109/TPAMI.2024.3387349
中图分类号
TP18 [人工智能理论];
学科分类号
081104 ; 0812 ; 0835 ; 1405 ;
摘要
Nearly all existing scene graph generation (SGG) models have overlooked the ground-truth annotation qualities of mainstream SGG datasets, i.e., they assume: 1) all the manually annotated positive samples are equally correct; 2) all the un-annotated negative samples are absolutely background. In this article, we argue that neither of the assumptions applies to SGG: there are numerous "noisy" ground-truth predicate labels that break these two assumptions and harm the training of unbiased SGG models. To this end, we propose a novel NoIsy label CorrEction and Sample Training strategy for SGG: NICEST, which rules out these noisy label issues by generating high-quality samples and designing an effective training strategy. Specifically, it consists of: 1) NICE: it detects noisy samples and then reassigns higher-quality soft predicate labels to them. To achieve this goal, NICE contains three main steps: negative Noisy Sample Detection (Neg-NSD), positive NSD (Pos-NSD), and Noisy Sample Correction (NSC). First, in Neg-NSD, it is treated as an out-of-distribution detection problem, and the pseudo labels are assigned to all detected noisy negative samples. Then, in Pos-NSD, we use a density-based clustering algorithm to detect noisy positive samples. Lastly, in NSC, we use weighted KNN to reassign more robust soft predicate labels rather than hard labels to all noisy positive samples. 2) NIST: it is a multi-teacher knowledge distillation based training strategy, which enables the model to learn unbiased fusion knowledge. A dynamic trade-off weighting strategy in NIST is designed to penalize the bias of different teachers. Due to the model-agnostic nature of both NICE and NIST, NICEST can be seamlessly incorporated into any SGG architecture to boost its performance on different predicate categories. In addition, to better assess the generalization ability of SGG models, we propose a new benchmark, VG-OOD, by reorganizing the prevalent VG dataset. This reorganization deliberately makes the predicate distributions between the training and test sets as different as possible for each subject-object category pair. This new benchmark helps disentangle the influence of subject-object category biases. Extensive ablations and results on different backbones and tasks have attested to the effectiveness and generalization ability of each component of NICEST.
引用
收藏
页码:6873 / 6888
页数:16
相关论文
共 92 条
[1]   Don't Just Assume; Look and Answer: Overcoming Priors for Visual Question Answering [J].
Agrawal, Aishwarya ;
Batra, Dhruv ;
Parikh, Devi ;
Kembhavi, Aniruddha .
2018 IEEE/CVF CONFERENCE ON COMPUTER VISION AND PATTERN RECOGNITION (CVPR), 2018, :4971-4980
[2]   A Comprehensive Survey of Scene Graphs: Generation and Application [J].
Chang, Xiaojun ;
Ren, Pengzhen ;
Xu, Pengfei ;
Li, Zhihui ;
Chen, Xiaojiang ;
Hauptmann, Alex .
IEEE TRANSACTIONS ON PATTERN ANALYSIS AND MACHINE INTELLIGENCE, 2023, 45 (01) :1-26
[3]   Addressing Predicate Overlap in Scene Graph Generation with Semantic Granularity Controller [J].
Chen, Guikun ;
Li, Lin ;
Luo, Yawei ;
Xiao, Jun .
2023 IEEE INTERNATIONAL CONFERENCE ON MULTIMEDIA AND EXPO, ICME, 2023, :78-83
[4]  
Chen GB, 2017, ADV NEUR IN, V30
[5]   Rethinking Data Augmentation for Robust Visual Question Answering [J].
Chen, Long ;
Zheng, Yuhang ;
Xiao, Jun .
COMPUTER VISION, ECCV 2022, PT XXXVI, 2022, 13696 :95-112
[6]  
Chen L, 2023, Arxiv, DOI arXiv:2110.01013
[7]   Counterfactual Samples Synthesizing for Robust Visual Question Answering [J].
Chen, Long ;
Yan, Xin ;
Xiao, Jun ;
Zhang, Hanwang ;
Pu, Shiliang ;
Zhuang, Yueting .
2020 IEEE/CVF CONFERENCE ON COMPUTER VISION AND PATTERN RECOGNITION (CVPR 2020), 2020, :10797-10806
[8]   Counterfactual Critic Multi-Agent Training for Scene Graph Generation [J].
Chen, Long ;
Zhang, Hanwang ;
Xiao, Jun ;
He, Xiangnan ;
Pu, Shiliang ;
Chang, Shih-Fu .
2019 IEEE/CVF INTERNATIONAL CONFERENCE ON COMPUTER VISION (ICCV 2019), 2019, :4612-4622
[9]   Knowledge-Embedded Routing Network for Scene Graph Generation [J].
Chen, Tianshui ;
Yu, Weihao ;
Chen, Riquan ;
Lin, Liang .
2019 IEEE/CVF CONFERENCE ON COMPUTER VISION AND PATTERN RECOGNITION (CVPR 2019), 2019, :6156-6164
[10]   Recovering the Unbiased Scene Graphs from the Biased Ones [J].
Chiou, Meng-Jiun ;
Ding, Henghui ;
Yan, Hanshu ;
Wang, Changhu ;
Zimmermann, Roger ;
Feng, Jiashi .
PROCEEDINGS OF THE 29TH ACM INTERNATIONAL CONFERENCE ON MULTIMEDIA, MM 2021, 2021, :1581-1590