DCAT: Dual Cross-Attention-Based Transformer for Change Detection

被引:16
作者
Zhou, Yuan [1 ,2 ]
Huo, Chunlei [1 ,2 ,3 ]
Zhu, Jiahang [1 ,2 ]
Huo, Leigang [4 ]
Pan, Chunhong [2 ]
机构
[1] Univ Chinese Acad Sci, Sch Artificial Intelligence, Beijing 101408, Peoples R China
[2] Chinese Acad Sci, Inst Automation, Natl Lab Pattern Recognit, Beijing 100190, Peoples R China
[3] Univ Sci & Technol Beijing, Sch Automation & Elect Engn, Beijing 100083, Peoples R China
[4] Nanning Normal Univ, Sch Comp & Informat Engn, Nanning 530001, Peoples R China
基金
中国国家自然科学基金;
关键词
change detection; transformer; dual cross-attention; remote sensing; IMAGE; CONNECTIONS; FEATURES; NETWORK;
D O I
10.3390/rs15092395
中图分类号
X [环境科学、安全科学];
学科分类号
08 ; 0830 ;
摘要
Several transformer-based methods for change detection (CD) in remote sensing images have been proposed, with Siamese-based methods showing promising results due to their two-stream feature extraction structure. However, these methods ignore the potential of the cross-attention mechanism to improve change feature discrimination and thus, may limit the final performance. Additionally, using either high-frequency-like fast change or low-frequency-like slow change alone may not effectively represent complex bi-temporal features. Given these limitations, we have developed a new approach that utilizes the dual cross-attention-transformer (DCAT) method. This method mimics the visual change observation procedure of human beings and interacts with and merges bi-temporal features. Unlike traditional Siamese-based CD frameworks, the proposed method extracts multi-scale features and models patch-wise change relationships by connecting a series of hierarchically structured dual cross-attention blocks (DCAB). DCAB is based on a hybrid dual branch mixer that combines convolution and transformer to extract and fuse local and global features. It calculates two types of cross-attention features to effectively learn comprehensive cues with both low- and high-frequency information input from paired CD images. This helps enhance discrimination between the changed and unchanged regions during feature extraction. The feature pyramid fusion network is more lightweight than the encoder and produces powerful multi-scale change representations by aggregating features from different layers. Experiments on four CD datasets demonstrate the advantages of DCAT architecture over other state-of-the-art methods.
引用
收藏
页数:30
相关论文
共 68 条
[11]  
Daudt RC, 2018, INT GEOSCI REMOTE SE, P2115, DOI 10.1109/IGARSS.2018.8518015
[12]  
Devlin J, 2019, Arxiv, DOI [arXiv:1810.04805, DOI 10.48550/ARXIV.1810.04805]
[13]  
Ding L, 2022, Arxiv, DOI [arXiv:2108.06103, 10.1109/TGRS.2022.3154390, DOI 10.1109/TGRS.2022.3154390]
[14]   DSA-Net: A novel deeply supervised attention-guided network for building change detection in high-resolution remote sensing images [J].
Ding, Qing ;
Shao, Zhenfeng ;
Huang, Xiao ;
Altan, Orhan .
INTERNATIONAL JOURNAL OF APPLIED EARTH OBSERVATION AND GEOINFORMATION, 2021, 105
[15]  
Dosovitskiy A, 2021, Arxiv, DOI arXiv:2010.11929
[16]   An improved change detection approach using tri-temporal logic-verified change vector analysis [J].
Du, Peijun ;
Wang, Xin ;
Chen, Dongmei ;
Liu, Sicong ;
Lin, Cong ;
Meng, Yaping .
ISPRS JOURNAL OF PHOTOGRAMMETRY AND REMOTE SENSING, 2020, 161 :278-293
[17]   Taming Transformers for High-Resolution Image Synthesis [J].
Esser, Patrick ;
Rombach, Robin ;
Ommer, Bjoern .
2021 IEEE/CVF CONFERENCE ON COMPUTER VISION AND PATTERN RECOGNITION, CVPR 2021, 2021, :12868-12878
[18]   SNUNet-CD: A Densely Connected Siamese Network for Change Detection of VHR Images [J].
Fang, Sheng ;
Li, Kaiyu ;
Shao, Jinyuan ;
Li, Zhe .
IEEE GEOSCIENCE AND REMOTE SENSING LETTERS, 2022, 19
[19]  
Fu L., 2022, arXiv, DOI DOI 10.48550/ARXIV.2212.03035
[20]  
Fujita A, 2017, PROCEEDINGS OF THE FIFTEENTH IAPR INTERNATIONAL CONFERENCE ON MACHINE VISION APPLICATIONS - MVA2017, P5, DOI 10.23919/MVA.2017.7986759