Shallow-Guided Transformer for Semantic Segmentation of Hyperspectral Remote Sensing Imagery

被引：10

作者：

Chen, Yuhan ^{[1
]}

Liu, Pengyuan ^{[2
]}

Zhao, Jiechen ^{[3
]}

Huang, Kaijian ^{[4
]}

Yan, Qingyun ^{[1
]}

机构：

[1] Nanjing Univ Informat Sci & Technol, Sch Remote Sensing & Geomatics Engn, Nanjing 210044, Peoples R China

[2] Nanjing Univ Informat Sci & Technol, Sch Geog Sci, Nanjing 210044, Peoples R China

[3] Harbin Engn Univ, Qingdao Innovat & Dev Base Ctr, Qingdao 266000, Peoples R China

[4] Huizhou Univ, Sch Elect Informat & Elect Engn, Huizhou 516007, Peoples R China

来源：

REMOTE SENSING | 2023年 / 15卷 / 13期

基金：

中国国家自然科学基金;

关键词：

vision transformer; convolutional neural networks (CNNs); feature representations; hyperspectral images (HSIs); semantic segmentation; NETWORK;

D O I：

10.3390/rs15133366

中图分类号：

X [环境科学、安全科学];

学科分类号：

08 ; 0830 ;

摘要：

Convolutional neural networks (CNNs) have achieved great progress in the classification of surface objects with hyperspectral data, but due to the limitations of convolutional operations, CNNs cannot effectively interact with contextual information. Transformer succeeds in solving this problem, and thus has been widely used to classify hyperspectral surface objects in recent years. However, the huge computational load of Transformer poses a challenge in hyperspectral semantic segmentation tasks. In addition, the use of single Transformer discards the local correlation, making it ineffective for remote sensing tasks with small datasets. Therefore, we propose a new Transformer layered architecture that combines Transformer with CNN, adopts a feature dimensionality reduction module and a Transformer-style CNN module to extract shallow features and construct texture constraints, and employs the original Transformer Encoder to extract deep features. Furthermore, we also designed a simple Decoder to process shallow spatial detail information and deep semantic features separately. Experimental results based on three publicly available hyperspectral datasets show that our proposed method has significant advantages compared with other traditional CNN, Transformer-type models.

引用

页数：23

共 49 条

[1] 3-D Deep Learning Approach for Remote Sensing Image Classification [J].

Ben Hamida, Amina ;

Benoit, Alexandre ;

Lambert, Patrick ;

Ben Amar, Chokri .

IEEE TRANSACTIONS ON GEOSCIENCE AND REMOTE SENSING, 2018, 56 (08) :4420-4434

[2] Mask-guided Spectral-wise Transformer for Efficient Hyperspectral Image Reconstruction [J].

Cai, Yuanhao ;

Lin, Jing ;

Hu, Xiaowan ;

Wang, Haoqian ;

Yuan, Xin ;

Zhang, Yulun ;

Timofte, Radu ;

Van Gool, Luc .

2022 IEEE/CVF CONFERENCE ON COMPUTER VISION AND PATTERN RECOGNITION (CVPR 2022), 2022, :17481-17490

[3] Making Vision Transformers Efficient from A Token Sparsification View [J].

Chang, Shuning ;

Wang, Pichao ;

Lin, Ming ;

Wang, Fan ;

Zhang, David Junhao ;

Jin, Rong ;

Shou, Mike Zheng .

2023 IEEE/CVF CONFERENCE ON COMPUTER VISION AND PATTERN RECOGNITION, CVPR, 2023, :6195-6205

[4]

Chaurasia A, 2017, 2017 IEEE VISUAL COMMUNICATIONS AND IMAGE PROCESSING (VCIP)

[5]

Chen LC, 2017, Arxiv, DOI arXiv:1706.05587

[6] Encoder-Decoder with Atrous Separable Convolution for Semantic Image Segmentation [J].

Chen, Liang-Chieh ;

Zhu, Yukun ;

Papandreou, George ;

Schroff, Florian ;

Adam, Hartwig .

COMPUTER VISION - ECCV 2018, PT VII, 2018, 11211 :833-851

[7] Hyperspectral Remote-Sensing Classification Combining Transformer and Multiscale Residual Mechanisms [J].

Chen Yuhan ;

Wang Bo ;

Yan Qingyun ;

Huang Bingjie ;

Jia Tong ;

Xue Bin .

LASER & OPTOELECTRONICS PROGRESS, 2023, 60 (12)

[8]

Chu XX, 2021, ADV NEUR IN

[9] ConViT: improving vision transformers with soft convolutional inductive biases [J].

d'Ascoli, Stephane ;

Touvron, Hugo ;

Leavitt, Matthew L. ;

Morcos, Ari S. ;

Biroli, Giulio ;

Sagun, Levent .

JOURNAL OF STATISTICAL MECHANICS-THEORY AND EXPERIMENT, 2022, 2022 (11)

[10]

Devlin J, 2019, Arxiv, DOI [arXiv:1810.04805, 10.48550/arxiv.1810.04805]

← 1 2 3 4 5 →