NICEST: Noisy Label Correction and Training for Robust Scene Graph Generation

被引:6
作者
Li, Lin [1 ]
Xiao, Jun [1 ]
Shi, Hanrong [1 ]
Zhang, Hanwang [2 ]
Yang, Yi [1 ]
Liu, Wei [3 ]
Chen, Long [4 ]
机构
[1] Zhejiang Univ, Coll Comp Sci, Hangzhou 310027, Peoples R China
[2] Nanyang Technol Univ, Sch Comp Sci & Engn, Singapore 639798, Singapore
[3] Tencent, Shenzhen 518000, Peoples R China
[4] Hong Kong Univ Sci & Technol, Dept Comp Sci & Engn, Clear Water Bay, Hong Kong, Peoples R China
基金
中国国家自然科学基金;
关键词
Noise measurement; Training; Visualization; Annotations; NIST; Benchmark testing; Task analysis; Multi-Teacher knowledge distillation; noisy label learning; out-of-distribution; scene graph generation;
D O I
10.1109/TPAMI.2024.3387349
中图分类号
TP18 [人工智能理论];
学科分类号
081104 ; 0812 ; 0835 ; 1405 ;
摘要
Nearly all existing scene graph generation (SGG) models have overlooked the ground-truth annotation qualities of mainstream SGG datasets, i.e., they assume: 1) all the manually annotated positive samples are equally correct; 2) all the un-annotated negative samples are absolutely background. In this article, we argue that neither of the assumptions applies to SGG: there are numerous "noisy" ground-truth predicate labels that break these two assumptions and harm the training of unbiased SGG models. To this end, we propose a novel NoIsy label CorrEction and Sample Training strategy for SGG: NICEST, which rules out these noisy label issues by generating high-quality samples and designing an effective training strategy. Specifically, it consists of: 1) NICE: it detects noisy samples and then reassigns higher-quality soft predicate labels to them. To achieve this goal, NICE contains three main steps: negative Noisy Sample Detection (Neg-NSD), positive NSD (Pos-NSD), and Noisy Sample Correction (NSC). First, in Neg-NSD, it is treated as an out-of-distribution detection problem, and the pseudo labels are assigned to all detected noisy negative samples. Then, in Pos-NSD, we use a density-based clustering algorithm to detect noisy positive samples. Lastly, in NSC, we use weighted KNN to reassign more robust soft predicate labels rather than hard labels to all noisy positive samples. 2) NIST: it is a multi-teacher knowledge distillation based training strategy, which enables the model to learn unbiased fusion knowledge. A dynamic trade-off weighting strategy in NIST is designed to penalize the bias of different teachers. Due to the model-agnostic nature of both NICE and NIST, NICEST can be seamlessly incorporated into any SGG architecture to boost its performance on different predicate categories. In addition, to better assess the generalization ability of SGG models, we propose a new benchmark, VG-OOD, by reorganizing the prevalent VG dataset. This reorganization deliberately makes the predicate distributions between the training and test sets as different as possible for each subject-object category pair. This new benchmark helps disentangle the influence of subject-object category biases. Extensive ablations and results on different backbones and tasks have attested to the effectiveness and generalization ability of each component of NICEST.
引用
收藏
页码:6873 / 6888
页数:16
相关论文
共 50 条
[21]   MLMG-SGG: Multilabel Scene Graph Generation With Multigrained Features [J].
Li, Xuewei ;
Miao, Peihan ;
Li, Songyuan ;
Li, Xi .
IEEE TRANSACTIONS ON IMAGE PROCESSING, 2024, 33 :1549-1559
[22]   Adaptive Fine-Grained Predicates Learning for Scene Graph Generation [J].
Lyu, Xinyu ;
Gao, Lianli ;
Zeng, Pengpeng ;
Shen, Heng Tao ;
Song, Jingkuan .
IEEE TRANSACTIONS ON PATTERN ANALYSIS AND MACHINE INTELLIGENCE, 2023, 45 (11) :13921-13940
[23]   Graph Regularized AutoFuse: Robust Sensor Fusion With Noisy Labels [J].
Sahu, Saurabh ;
Kumar, Kriti ;
Majumdar, Angshul ;
Kumar, A. Anil ;
Chandra, M. Girish .
IEEE SENSORS LETTERS, 2025, 9 (02)
[24]   BiFormer for Scene Graph Generation Based on VisionNet With Taylor Hiking Optimization Algorithm [J].
Monesh, S. ;
Senthilkumar, N. C. .
IEEE ACCESS, 2025, 13 :57207-57222
[25]   Dual-Branch Hybrid Learning Network for Unbiased Scene Graph Generation [J].
Zheng, Chaofan ;
Gao, Lianli ;
Lyu, Xinyu ;
Zeng, Pengpeng ;
El Saddik, Abdulmotaleb ;
Shen, Heng Tao .
IEEE TRANSACTIONS ON CIRCUITS AND SYSTEMS FOR VIDEO TECHNOLOGY, 2024, 34 (03) :1743-1756
[26]   Spatial–Temporal Knowledge-Embedded Transformer for Video Scene Graph Generation [J].
Pu, Tao ;
Chen, Tianshui ;
Wu, Hefeng ;
Lu, Yongyi ;
Lin, Liang .
IEEE TRANSACTIONS ON IMAGE PROCESSING, 2024, 33 :556-568
[27]   Beware of Overcorrection: Scene-induced Commonsense Graph for Scene Graph Generation [J].
Chen, Lianggangxu ;
Lu, Jiale ;
Song, Youqi ;
Wang, Changbo ;
He, Gaoqi .
PROCEEDINGS OF THE 31ST ACM INTERNATIONAL CONFERENCE ON MULTIMEDIA, MM 2023, 2023, :2888-2897
[28]   Toward a Unified Transformer-Based Framework for Scene Graph Generation and Human-Object Interaction Detection [J].
He, Tao ;
Gao, Lianli ;
Song, Jingkuan ;
Li, Yuan-Fang .
IEEE TRANSACTIONS ON IMAGE PROCESSING, 2023, 32 :6274-6288
[29]   Robust Mechanical Fault Diagnosis With Noisy Label Based on Multistage True Label Distribution Learning [J].
Wang, Huan ;
Li, Yan-Fu .
IEEE TRANSACTIONS ON RELIABILITY, 2023, 72 (03) :975-988
[30]   Multi-Label Meta Weighting for Long-Tailed Dynamic Scene Graph Generation [J].
Chen, Shuo ;
Du, Yingjun ;
Mettes, Pascal ;
Snoek, Cees G. M. .
PROCEEDINGS OF THE 2023 ACM INTERNATIONAL CONFERENCE ON MULTIMEDIA RETRIEVAL, ICMR 2023, 2023, :39-47