RF-Next: Efficient Receptive Field Search for Convolutional Neural Networks

被引：22

作者：

Gao, Shanghua ^{[1
]}

Li, Zhong-Yu ^{[1
]}

Han, Qi ^{[1
]}

Cheng, Ming-Ming ^{[1
]}

Wang, Liang ^{[2
]}

机构：

[1] Nankai Univ, TMCC, CS, Tianjin 300350, Peoples R China

[2] Natl Lab Pattern Recognit, Beijing 100190, Peoples R China

来源：

IEEE TRANSACTIONS ON PATTERN ANALYSIS AND MACHINE INTELLIGENCE | 2023年 / 45卷 / 03期

关键词：

Dilation; receptive field; spatial convolutional network; temporal convolutional network; temporal action segmentation;

D O I：

10.1109/TPAMI.2022.3183829

中图分类号：

TP18 [人工智能理论];

学科分类号：

081104 ; 0812 ; 0835 ; 1405 ;

摘要：

Temporal/spatial receptive fields of models play an important role in sequential/spatial tasks. Large receptive fields facilitate long-term relations, while small receptive fields help to capture the local details. Existing methods construct models with hand-designed receptive fields in layers. Can we effectively search for receptive field combinations to replace hand-designed patterns? To answer this question, we propose to find better receptive field combinations through a global-to-local search scheme. Our search scheme exploits both global search to find the coarse combinations and local search to get the refined receptive field combinations further. The global search finds possible coarse combinations other than human-designed patterns. On top of the global search, we propose an expectation-guided iterative local search scheme to refine combinations effectively. Our RF-Next models, plugging receptive field search to various models, boost the performance on many tasks, e.g., temporal action segmentation, object detection, instance segmentation, and speech synthesis.

引用

页码：2984 / 3002

页数：19

共 154 条

[1] MS-TCN: Multi-Stage Temporal Convolutional Network for Action Segmentation
Abu Farha, Yazan
Gall, Juergen
[J]. 2019 IEEE/CVF CONFERENCE ON COMPUTER VISION AND PATTERN RECOGNITION (CVPR 2019), 2019, : 3570 - 3579
[2] Arik SO, 2017, PR MACH LEARN RES, V70
[3] Bai SJ, 2018, Arxiv, DOI [arXiv:1803.01271, 10.48550/arXiv.1803.01271]
[4] Recognition of Complex Events: Exploiting Temporal Dynamics between Underlying Concepts
Bhattacharya, Subhabrata
Kalaych, Mahdi M.
Sukthankar, Rahul
Shah, Mubarak
[J]. 2014 IEEE CONFERENCE ON COMPUTER VISION AND PATTERN RECOGNITION (CVPR), 2014, : 2243 - 2250
[5] Cascade R-CNN: High Quality Object Detection and Instance Segmentation
Cai, Zhaowei
Vasconcelos, Nuno
[J]. IEEE TRANSACTIONS ON PATTERN ANALYSIS AND MACHINE INTELLIGENCE, 2021, 43 (05) : 1483 - 1498
[6] Quo Vadis, Action Recognition? A New Model and the Kinetics Dataset
Carreira, Joao
Zisserman, Andrew
[J]. 30TH IEEE CONFERENCE ON COMPUTER VISION AND PATTERN RECOGNITION (CVPR 2017), 2017, : 4724 - 4733
[7] Hybrid Task Cascade for Instance Segmentation
Chen, Kai
Pang, Jiangmiao
Wang, Jiaqi
Xiong, Yu
Li, Xiaoxiao
Sun, Shuyang
Feng, Wansen
Liu, Ziwei
Shi, Jianping
Ouyang, Wanli
Loy, Chen Change
Lin, Dahua
[J]. 2019 IEEE/CVF CONFERENCE ON COMPUTER VISION AND PATTERN RECOGNITION (CVPR 2019), 2019, : 4969 - 4978
[8] Chen LC, 2017, Arxiv, DOI arXiv:1706.05587
[9] DeepLab: Semantic Image Segmentation with Deep Convolutional Nets, Atrous Convolution, and Fully Connected CRFs
Chen, Liang-Chieh
Papandreou, George
Kokkinos, Iasonas
Murphy, Kevin
Yuille, Alan L.
[J]. IEEE TRANSACTIONS ON PATTERN ANALYSIS AND MACHINE INTELLIGENCE, 2018, 40 (04) : 834 - 848
[10] Chen MH, 2020, PROC CVPR IEEE, P9451, DOI 10.1109/CVPR42600.2020.00947

← 1 2 3 4 5 6 7 8 9 10 →