Improving RGB-D Salient Object Detection via Modality-Aware Decoder

被引:33
|
作者
Song, Mengke [1 ,2 ]
Song, Wenfeng [3 ]
Yang, Guowei [4 ]
Chen, Chenglizhao [1 ,2 ]
机构
[1] China Univ Petr East China, Coll Comp Sci & Technol, Qingdao 266580, Peoples R China
[2] China Univ Petr East China, Qingdao Inst Software, Qingdao 266580, Peoples R China
[3] Beijing Informat Sci & Technol Univ, Comp Sch, Beijing 100192, Peoples R China
[4] Qingdao Univ, Sch Elect Informat, Qingdao 266071, Peoples R China
基金
中国国家自然科学基金;
关键词
Decoding; Object detection; Training; Task analysis; Saliency detection; Image segmentation; Feature extraction; RGB-D salient object detection; modality-aware fusion; deep learning; GRAPH CONVOLUTION NETWORK; IMAGE; ATTENTION; FUSION;
D O I
10.1109/TIP.2022.3205747
中图分类号
TP18 [人工智能理论];
学科分类号
081104 ; 0812 ; 0835 ; 1405 ;
摘要
Most existing RGB-D salient object detection (SOD) methods are primarily focusing on cross-modal and cross-level saliency fusion, which has been proved to be efficient and effective. However, these methods still have a critical limitation, i.e., their fusion patterns - typically the combination of selective characteristics and its variations, are too highly dependent on the network's non-linear adaptability. In such methods, the balances between RGB and D (Depth) are formulated individually considering the intermediate feature slices, but the relation at the modality level may not be learned properly. The optimal RGB-D combinations differ depending on the RGB-D scenarios, and the exact complementary status is frequently determined by multiple modality-level factors, such as D quality, the complexity of the RGB scene, and degree of harmony between them. Therefore, given the existing approaches, it may be difficult for them to achieve further performance breakthroughs, as their methodologies belong to some methods that are somewhat less modality sensitive. To conquer this problem, this paper presents the Modality-aware Decoder (MaD). The critical technical innovations include a series of feature embedding, modality reasoning, and feature back-projecting and collecting strategies, all of which upgrade the widely-used multi-scale and multi-level decoding process to be modality-aware. Our MaD achieves competitive performance over other state-of-the-art (SOTA) models without using any fancy tricks in the decoder's design. Codes and results will be publicly available at https://github.com/MengkeSong/MaD.
引用
收藏
页码:6124 / 6138
页数:15
相关论文
共 50 条
  • [1] Dynamic Selective Network for RGB-D Salient Object Detection
    Wen, Hongfa
    Yan, Chenggang
    Zhou, Xiaofei
    Cong, Runmin
    Sun, Yaoqi
    Zheng, Bolun
    Zhang, Jiyong
    Bao, Yongjun
    Ding, Guiguang
    IEEE TRANSACTIONS ON IMAGE PROCESSING, 2021, 30 : 9179 - 9192
  • [2] EM-Trans: Edge-Aware Multimodal Transformer for RGB-D Salient Object Detection
    Chen, Geng
    Wang, Qingyue
    Dong, Bo
    Ma, Ruitao
    Liu, Nian
    Fu, Huazhu
    Xia, Yong
    IEEE TRANSACTIONS ON NEURAL NETWORKS AND LEARNING SYSTEMS, 2025, 36 (02) : 3175 - 3188
  • [3] Hierarchical Alternate Interaction Network for RGB-D Salient Object Detection
    Li, Gongyang
    Liu, Zhi
    Chen, Minyu
    Bai, Zhen
    Lin, Weisi
    Ling, Haibin
    IEEE TRANSACTIONS ON IMAGE PROCESSING, 2021, 30 : 3528 - 3542
  • [4] 3-D Convolutional Neural Networks for RGB-D Salient Object Detection and Beyond
    Chen, Qian
    Zhang, Zhenxi
    Lu, Yanye
    Fu, Keren
    Zhao, Qijun
    IEEE TRANSACTIONS ON NEURAL NETWORKS AND LEARNING SYSTEMS, 2024, 35 (03) : 4309 - 4323
  • [5] RGB-D Salient Object Detection Using Saliency and Edge Reverse Attention
    Ikeda, Tomoki
    Ikehara, Masaaki
    IEEE ACCESS, 2023, 11 : 68818 - 68825
  • [6] LIANet: Layer Interactive Attention Network for RGB-D Salient Object Detection
    Han, Yibo
    Wang, Liejun
    Du, Anyu
    Jiang, Shaochen
    IEEE ACCESS, 2022, 10 : 25435 - 25447
  • [7] FCMNet: Frequency-aware cross-modality attention networks for RGB-D salient object detection
    Jin, Xiao
    Guo, Chunle
    He, Zhen
    Xu, Jing
    Wang, Yongwei
    Su, Yuting
    NEUROCOMPUTING, 2022, 491 : 414 - 425
  • [8] CDNet: Complementary Depth Network for RGB-D Salient Object Detection
    Jin, Wen-Da
    Xu, Jun
    Han, Qi
    Zhang, Yi
    Cheng, Ming-Ming
    IEEE TRANSACTIONS ON IMAGE PROCESSING, 2021, 30 : 3376 - 3390
  • [9] DGFNet: Depth-Guided Cross-Modality Fusion Network for RGB-D Salient Object Detection
    Xiao, Fen
    Pu, Zhengdong
    Chen, Jiaqi
    Gao, Xieping
    IEEE TRANSACTIONS ON MULTIMEDIA, 2024, 26 : 2648 - 2658
  • [10] Employing Bilinear Fusion and Saliency Prior Information for RGB-D Salient Object Detection
    Huang, Nianchang
    Yang, Yang
    Zhang, Dingwen
    Zhang, Qiang
    Han, Jungong
    IEEE TRANSACTIONS ON MULTIMEDIA, 2022, 24 : 1651 - 1664