Consistency-constrained RGB-T crowd counting via mutual information maximization

被引：3

作者：

Guo, Qiang ^{[1
]}

Yuan, Pengcheng ^{[1
]}

Huang, Xiangming ^{[1
]}

Ye, Yangdong ^{[1
]}

机构：

[1] Zhengzhou Univ, Sch Comp & Artificial Intelligence, 100 Sci Ave, Zhengzhou 450001, Henan, Peoples R China

来源：

COMPLEX & INTELLIGENT SYSTEMS | 2024年 / 10卷 / 04期

基金：

中国国家自然科学基金;

关键词：

Cross-modal; RGB-T crowd counting; Consistent information; Mutual information;

D O I：

10.1007/s40747-024-01427-x

中图分类号：

TP18 [人工智能理论];

学科分类号：

081104 ; 0812 ; 0835 ; 1405 ;

摘要：

The incorporation of thermal imaging data in RGB-T images has demonstrated its usefulness in cross-modal crowd counting by offering complementary information to RGB representations. Despite achieving satisfactory results in RGB-T crowd counting, many existing methods still face two significant limitations: (1) The oversight of the heterogeneous gap between modalities complicates the effective integration of multimodal features. (2) The absence of mining consistency hinders the full exploitation of the unique complementary strengths inherent in each modality. To this end, we present C4-MIM, a novel Consistency-constrained RGB-T Crowd Counting approach via Mutual Information Maximization. It effectively leverages multimodal information by learning the consistency between the RGB and thermal modalities, thereby enhancing the performance of cross-modal counting. Specifically, we first advocate extracting feature representations of different modalities in a shared encoder to moderate the heterogeneous gap since they obey the identical coding rules with shared parameters. Then, we intend to mine the consistent information of different modalities to better learn conducive information and improve the performance of feature representations. To this end, we formulate the complementarity of multimodality representations as a mutual information maximization regularizer to maximize the consistent information of different modalities, in which the consistency would be maximally attained before combining the multimodal information. Finally, we simply aggregate the feature representations of the different modalities and send them into a regressor to output the density maps. The proposed approach can be implemented by arbitrary backbone networks and is quite robust in the face of single modality unavailable or serious compromised. Extensively experiments have been conducted on the RGBT-CC and DroneRGBT benchmarks to evaluate the effectiveness and robustness of the proposed approach, demonstrating its superior performance compared to the SOTA approaches.

引用

页码：5049 / 5070

页数：22

共 54 条

[11] Multimodal medical image fusion with convolution sparse representation and mutual information correlation in NSST domain [J].

Guo, Peng ;

Xie, Guoqi ;

Li, Renfa ;

Hu, Hui .

COMPLEX & INTELLIGENT SYSTEMS, 2023, 9 (01) :317-328

[12] Learning a deep network with cross-hierarchy aggregation for crowd counting [J].

Guo, Qiang ;

Zeng, Xin ;

Hu, Shizhe ;

Phoummixay, Sonephet ;

Ye, Yangdong .

KNOWLEDGE-BASED SYSTEMS, 2021, 213

[13] Multi-Source Multi-Scale Counting in Extremely Dense Crowd Images [J].

Idrees, Haroon ;

Saleemi, Imran ;

Seibert, Cody ;

Shah, Mubarak .

2013 IEEE CONFERENCE ON COMPUTER VISION AND PATTERN RECOGNITION (CVPR), 2013, :2547-2554

[14] Beyond Counting: Comparisons of Density Maps for Crowd Analysis Tasks-Counting, Detection, and Tracking [J].

Kang, Di ;

Ma, Zheng ;

Chan, Antoni B. .

IEEE TRANSACTIONS ON CIRCUITS AND SYSTEMS FOR VIDEO TECHNOLOGY, 2019, 29 (05) :1408-1422

[15] RankMI: A Mutual Information Maximizing Ranking Loss [J].

Kemertas, Mete ;

Pishdad, Leila ;

Derpanis, Konstantinos G. ;

Fazly, Afsaneh .

2020 IEEE/CVF CONFERENCE ON COMPUTER VISION AND PATTERN RECOGNITION (CVPR 2020), 2020, :14350-14359

[16] Lowering glycemic levels via gastrointestinal tract factors: the roles of dietary fiber, polyphenols, and their combination [J].

Li, Fuhua ;

Zeng, Kaifang ;

Ming, Jian .

CRITICAL REVIEWS IN FOOD SCIENCE AND NUTRITION, 2025, 65 (03) :575-611

[17] CSA-Net: Cross-modal scale-aware attention-aggregated network for RGB-T crowd counting [J].

Li, He ;

Zhang, Junge ;

Kong, Weihang ;

Shen, Jienan ;

Shao, Yuguang .

EXPERT SYSTEMS WITH APPLICATIONS, 2023, 213

[18] Learning the cross-modal discriminative feature representation for RGB-T crowd counting [J].

Li, He ;

Zhang, Shihui ;

Kong, Weihang .

KNOWLEDGE-BASED SYSTEMS, 2022, 257

[19] CSRNet: Dilated Convolutional Neural Networks for Understanding the Highly Congested Scenes [J].

Li, Yuhong ;

Zhang, Xiaofan ;

Chen, Deming .

2018 IEEE/CVF CONFERENCE ON COMPUTER VISION AND PATTERN RECOGNITION (CVPR), 2018, :1091-1100

[20] Consensus Graph Learning for Multi-View Clustering [J].

Li, Zhenglai ;

Tang, Chang ;

Liu, Xinwang ;

Zheng, Xiao ;

Zhang, Wei ;

Zhu, En .

IEEE TRANSACTIONS ON MULTIMEDIA, 2022, 24 :2461-2472

← 1 2 3 4 5 6 →