Hierarchical collaboration for referring image segmentation

被引：1

作者：

Zhang, Wei ^{[1
,2
]}

Cheng, Zesen ^{[3
]}

Chen, Jie ^{[2
,3
]}

Gao, Wen ^{[1
,2
]}

机构：

[1] Harbin Inst Technol, Sch Comp Sci & Technol, Shenzhen 518055, Peoples R China

[2] Peng Cheng Lab, Shenzhen 518000, Peoples R China

[3] Peking Univ, Sch Elect & Comp Engn, Shenzhen 518055, Peoples R China

来源：

NEUROCOMPUTING | 2025年 / 613卷

基金：

国家重点研发计划;

关键词：

Referring image segmentation; Image understanding; Cross-modal; TRANSFORMER; QUERY;

D O I：

10.1016/j.neucom.2024.128632

中图分类号：

TP18 [人工智能理论];

学科分类号：

081104 ; 0812 ; 0835 ; 1405 ;

摘要：

In the field of referring segmentation, top-down methods and bottom-up methods are the two prevailing approaches. Both of these methods inevitably exhibit certain drawbacks. Top-down methods are susceptible to Polar Negative (PN) errors due to their limited understanding of multi-modal fine-grained features. Bottom-up methods lack macro-level object positional information, making them susceptible to Inferior Positive (IP) errors. However, we find that the two approaches are highly complementary in addressing their respective weaknesses, but combining them directly through a simple average does not yield complementary advantages. Therefore, we proposed a hierarchical collaboration approach to explore the complementary characteristics of the existing two methods from the perspectives of fusion and interaction, aiming to achieve more precise segmentation results. We proposed the Complementary Feature Interaction (CFI) module, which enables top-down methods to access fine-grained information and allows bottom-up approaches to obtain object positional information interactively. Regarding integration, Gaussian Scoring Integration (GSI) models the Gaussian performance distributions of two branches and performs weighted integration by sampling confidence scores from these distributions. We integrate various top-down and bottom-up methods within the proposed architecture and conduct experiments on three standard datasets. The experimental results demonstrate that our method outperforms the state-of-theart independent segmentation algorithms. On the RefCOCO validation, test A and test B datasets, our proposed method achieved IoU scores of 77.51, 79.12, and 72.79, respectively. Extensive experiments demonstrate that our method can significantly improve segmentation accuracy when fusing different sub-methods.

引用

页数：13

共 82 条

[31] DynaMask: Dynamic Mask Selection for Instance Segmentation [J].

Li, Ruihuang ;

He, Chenhang ;

Li, Shuai ;

Zhang, Yabin ;

Zhang, Lei .

2023 IEEE/CVF CONFERENCE ON COMPUTER VISION AND PATTERN RECOGNITION (CVPR), 2023, :11279-11288

[32] Referring Image Segmentation via Recurrent Refinement Networks [J].

Li, Ruiyu ;

Li, Kaican ;

Kuo, Yi-Chun ;

Shu, Michelle ;

Qi, Xiaojuan ;

Shen, Xiaoyong ;

Jia, Jiaya .

2018 IEEE/CVF CONFERENCE ON COMPUTER VISION AND PATTERN RECOGNITION (CVPR), 2018, :5745-5753

[33] Logic-induced Diagnostic Reasoning for Semi-supervised Semantic Segmentation [J].

Liang, Chen ;

Wang, Wenguan ;

Miao, Jiaxu ;

Yang, Yi .

2023 IEEE/CVF INTERNATIONAL CONFERENCE ON COMPUTER VISION (ICCV 2023), 2023, :16151-16162

[34] Local-Global Context Aware Transformer for Language-Guided Video Segmentation [J].

Liang, Chen ;

Wang, Wenguan ;

Zhou, Tianfei ;

Miao, Jiaxu ;

Luo, Yawei ;

Yang, Yi .

IEEE TRANSACTIONS ON PATTERN ANALYSIS AND MACHINE INTELLIGENCE, 2023, 45 (08) :10055-10069

[35] Structured Attention Network for Referring Image Segmentation [J].

Lin, Liang ;

Yan, Pengxiang ;

Xu, Xiaoqian ;

Yang, Sibei ;

Zeng, Kun ;

Li, Guanbin .

IEEE TRANSACTIONS ON MULTIMEDIA, 2022, 24 :1922-1932

[36] Multi-Modal Mutual Attention and Iterative Interaction for Referring Image Segmentation [J].

Liu, Chang ;

Ding, Henghui ;

Zhang, Yulun ;

Jiang, Xudong .

IEEE TRANSACTIONS ON IMAGE PROCESSING, 2023, 32 :3054-3065

[37] Learning to Assemble Neural Module Tree Networks for Visual Grounding [J].

Liu, Daqing ;

Zhang, Hanwang ;

Wu, Feng ;

Zha, Zheng-Jun .

2019 IEEE/CVF INTERNATIONAL CONFERENCE ON COMPUTER VISION (ICCV 2019), 2019, :4672-4681

[38] Local-global coordination with transformers for referring image segmentation [J].

Liu, Fang ;

Kong, Yuqiu ;

Zhang, Lihe ;

Feng, Guang ;

Yin, Baocai .

NEUROCOMPUTING, 2023, 522 :39-52

[39] PolyFormer: Referring Image Segmentation as Sequential Polygon Generation [J].

Liu, Jiang ;

Ding, Hui ;

Cai, Zhaowei ;

Zhang, Yuting ;

Satzoda, Ravi Kumar ;

Mahadevan, Vijay ;

Manmatha, R. .

2023 IEEE/CVF CONFERENCE ON COMPUTER VISION AND PATTERN RECOGNITION (CVPR), 2023, :18653-18663

[40] Cross-Modal Progressive Comprehension for Referring Segmentation [J].

Liu, Si ;

Hui, Tianrui ;

Huang, Shaofei ;

Wei, Yunchao ;

Li, Bo ;

Li, Guanbin .

IEEE TRANSACTIONS ON PATTERN ANALYSIS AND MACHINE INTELLIGENCE, 2022, 44 (09) :4761-4775

← 1 2 3 4 5 6 7 8 9 →