A comparative attention framework for better few-shot object detection on aerial images

被引:3
作者
Le Jeune, Pierre [1 ,2 ]
Bahaduri, Bissmella [1 ]
Mokraoui, Anissa [1 ]
机构
[1] Univ Sorbonne Paris Nord, L2TI, 99 Ave Jean Baptiste Clement, F-93430 Villetaneuse, France
[2] COSE, 5 bis route St Leu, F-95360 Montmagny, France
关键词
Few-shot learning; Object detection; Aerial image; Few-shot object detection; Attention mechanisms;
D O I
10.1016/j.patcog.2024.111243
中图分类号
TP18 [人工智能理论];
学科分类号
081104 ; 0812 ; 0835 ; 1405 ;
摘要
Few-Shot Object Detection (FSOD) methods are mainly designed and evaluated on natural image datasets such as Pascal VOC and MS COCO. However, it is not clear whether the best methods for natural images are also the best for aerial images. Furthermore, a direct comparison of performance between FSOD methods difficult due to the wide variety of detection frameworks and training strategies. To this end, our contributions are twofold. First, we propose a benchmarking framework that provides a flexible environment to implement and compare attention-based FSOD methods. The proposed framework focuses on attention mechanisms and is divided into three modules: spatial alignment, global attention, and fusion layer. To remain competitive with existing methods, which often leverage complex training, we propose new augmentation techniques designed specifically for object detection. Using this framework, several FSOD methods are reimplemented and compared. This comparison highlights two distinct performance regimes on aerial and natural images: FSOD performs worse on aerial images. Our experiments confirm that small objects account for the poor performance. Small objects are difficult to detect, however in the few-shot regime, this challenge is largely reinforced. While the small object detection issue is well-known, to our knowledge this few-shot complication has never been reported in the literature. Second, always within the proposed framework, we develop a novel alignment method called Cross-Scales Query-Support Alignment (XQSA) for FSOD, to improve the detection of small objects. XQSA significantly outperforms the state-of-the-art on DOTA and DIOR, two aerial image datasets.
引用
收藏
页数:15
相关论文
共 40 条
[31]   Multi-scale structural kernel representation for object detection [J].
Wang, Hao ;
Wang, Qilong ;
Li, Peihua ;
Zuo, Wangmeng .
PATTERN RECOGNITION, 2021, 110
[32]  
Wang X., 2020, INT C MACH LEARN
[33]   Non-local Neural Networks [J].
Wang, Xiaolong ;
Girshick, Ross ;
Gupta, Abhinav ;
He, Kaiming .
2018 IEEE/CVF CONFERENCE ON COMPUTER VISION AND PATTERN RECOGNITION (CVPR), 2018, :7794-7803
[34]   DOTA: A Large-scale Dataset for Object Detection in Aerial Images [J].
Xia, Gui-Song ;
Bai, Xiang ;
Ding, Jian ;
Zhu, Zhen ;
Belongie, Serge ;
Luo, Jiebo ;
Datcu, Mihai ;
Pelillo, Marcello ;
Zhang, Liangpei .
2018 IEEE/CVF CONFERENCE ON COMPUTER VISION AND PATTERN RECOGNITION (CVPR), 2018, :3974-3983
[35]   Few-Shot Object Detection and Viewpoint Estimation for Objects in the Wild [J].
Xiao, Yang ;
Marlet, Renaud .
COMPUTER VISION - ECCV 2020, PT XVII, 2020, 12362 :192-210
[36]   Few-Shot Object Detection With Self-Adaptive Attention Network for Remote Sensing Images [J].
Xiao, Zixuan ;
Qi, Jiahao ;
Xue, Wei ;
Zhong, Ping .
IEEE JOURNAL OF SELECTED TOPICS IN APPLIED EARTH OBSERVATIONS AND REMOTE SENSING, 2021, 14 :4854-4865
[37]   Meta R-CNN : Towards General Solver for Instance-level Low-shot Learning [J].
Yan, Xiaopeng ;
Chen, Ziliang ;
Xu, Anni ;
Wang, Xiaoxi ;
Liang, Xiaodan ;
Lin, Liang .
2019 IEEE/CVF INTERNATIONAL CONFERENCE ON COMPUTER VISION (ICCV 2019), 2019, :9576-9585
[38]   Meta-DETR: Image-Level Few-Shot Detection With Inter-Class Correlation Exploitation [J].
Zhang, Gongjie ;
Luo, Zhipeng ;
Cui, Kaiwen ;
Lu, Shijian ;
Xing, Eric P. .
IEEE TRANSACTIONS ON PATTERN ANALYSIS AND MACHINE INTELLIGENCE, 2023, 45 (11) :12832-12843
[39]   Attention and boundary guided salient object detection [J].
Zhang, Qing ;
Shi, Yanjiao ;
Zhang, Xueqin .
PATTERN RECOGNITION, 2020, 107
[40]   Hallucination Improves Few-Shot Object Detection [J].
Zhang, Weilin ;
Wang, Yu-Xiong .
2021 IEEE/CVF CONFERENCE ON COMPUTER VISION AND PATTERN RECOGNITION, CVPR 2021, 2021, :13003-13012