PETNet: A YOLO-based prior enhanced transformer network for aerial image detection

被引:14
作者
Wang, Tianyu [1 ]
Ma, Zhongjing [1 ]
Yang, Tao [1 ]
Zou, Suli [1 ]
机构
[1] Beijing Inst Technol, Sch Automat, Beijing 100081, Peoples R China
基金
中国国家自然科学基金;
关键词
Deep learning; Transformer; Small object detection; Aerial image; OBJECT DETECTION; UAV;
D O I
10.1016/j.neucom.2023.126384
中图分类号
TP18 [人工智能理论];
学科分类号
081104 ; 0812 ; 0835 ; 1405 ;
摘要
Unmanned aerial vehicles (UAVs) have been applied to inspect in various scenarios due to their high effi-ciency, low cost, and excellent mobility. However, the objects in aerial images are much smaller and den-ser than general objects, causing it difficult for current object detection methods to achieve the expected results. To solve this issue, a prior enhanced Transformer network (PETNet) based on YOLO is proposed in this paper. Specifically, a novel prior enhanced Transformer (PET) module and a one-to-many feature fusion (OMFF) mechanism are proposed to embed into the network. Two additional detection heads are added to the shallow feature maps. In this work, PET is used to capture enhanced global information to improve the expressive ability of the network. The OMFF aims to fuse multi-type features to minimize the information loss of small objects. In addition, the added detection heads provide more possibility of detecting smaller-scale objects, and the extended multi-head parallel detection is more suitable for the multi-scale transformation of objects in aerial images. On the VisDrone-2021 and UAVDT databases, the proposed PETNet achieves state-of-the-art results with average precision (AP) of 35.3 and 21.5, respectively, which indicates that the proposed network is more suitable for aerial image detection and is of a great reference value.& COPY; 2023 Elsevier B.V. All rights reserved.
引用
收藏
页数:13
相关论文
共 52 条
  • [1] Bochkovskiy A, 2020, Arxiv, DOI arXiv:2004.10934
  • [2] Cascade R-CNN: Delving into High Quality Object Detection
    Cai, Zhaowei
    Vasconcelos, Nuno
    [J]. 2018 IEEE/CVF CONFERENCE ON COMPUTER VISION AND PATTERN RECOGNITION (CVPR), 2018, : 6154 - 6162
  • [3] Anchor-Free Oriented Proposal Generator for Object Detection
    Cheng, Gong
    Wang, Jiabao
    Li, Ke
    Xie, Xingxing
    Lang, Chunbo
    Yao, Yanqing
    Han, Junwei
    [J]. IEEE TRANSACTIONS ON GEOSCIENCE AND REMOTE SENSING, 2022, 60
  • [4] Dual-Aligned Oriented Detector
    Cheng, Gong
    Yao, Yanqing
    Li, Shengyang
    Li, Ke
    Xie, Xingxing
    Wang, Jiabao
    Yao, Xiwen
    Han, Junwei
    [J]. IEEE TRANSACTIONS ON GEOSCIENCE AND REMOTE SENSING, 2022, 60
  • [5] Chu XX, 2021, Arxiv, DOI arXiv:2102.10882
  • [6] Dosovitskiy A, 2021, Arxiv, DOI arXiv:2010.11929
  • [7] The Unmanned Aerial Vehicle Benchmark: Object Detection and Tracking
    Du, Dawei
    Qi, Yuankai
    Yu, Hongyang
    Yang, Yifan
    Duan, Kaiwen
    Li, Guorong
    Zhang, Weigang
    Huang, Qingming
    Tian, Qi
    [J]. COMPUTER VISION - ECCV 2018, PT X, 2018, 11214 : 375 - 391
  • [8] Coarse-grained Density Map Guided Object Detection in Aerial Images
    Duan, Chengzhen
    Wei, Zhiwei
    Zhang, Chi
    Qu, Siying
    Wang, Hongpeng
    [J]. 2021 IEEE/CVF INTERNATIONAL CONFERENCE ON COMPUTER VISION WORKSHOPS (ICCVW 2021), 2021, : 2789 - 2798
  • [9] Howard AG, 2017, Arxiv, DOI [arXiv:1704.04861, DOI 10.48550/ARXIV.1704.04861, 10.48550/arXiv.1704.04861]
  • [10] LeViT: a Vision Transformer in ConvNet's Clothing for Faster Inference
    Graham, Ben
    El-Nouby, Alaaeldin
    Touvron, Hugo
    Stock, Pierre
    Joulin, Armand
    Jegou, Herve
    Douze, Matthijs
    [J]. 2021 IEEE/CVF INTERNATIONAL CONFERENCE ON COMPUTER VISION (ICCV 2021), 2021, : 12239 - 12249