F-E3D: FPGA-based Acceleration of an Efficient 3D Convolutional Neural Network for Human Action Recognition

被引：37

作者：

Fan, Hongxiang ^{[1
]}

Luo, Cheng ^{[2
]}

Zeng, Chenglong ^{[3
]}

Ferianc, Martin ^{[1
]}

Que, Zhiqiang ^{[1
]}

Liu, Shuanglong ^{[1
]}

Niu, Xinyu ^{[4
]}

Luk, Wayne ^{[1
]}

机构：

[1] Imperial Coll London, Sch Engn, Dept Comp, London, England

[2] Fudan Univ, State Key Lab ASIC & Syst, Shanghai, Peoples R China

[3] Tianjin Univ, Sch Microelect, Tianjin, Peoples R China

[4] Corerain Technol Ltd, Shenzhen, Peoples R China

来源：

2019 IEEE 30TH INTERNATIONAL CONFERENCE ON APPLICATION-SPECIFIC SYSTEMS, ARCHITECTURES AND PROCESSORS (ASAP 2019) | 2019年

基金：

英国工程与自然科学研究理事会;

关键词：

D O I：

10.1109/ASAP.2019.00-44

中图分类号：

TP3 [计算技术、计算机技术];

学科分类号：

0812 ;

摘要：

Three-dimensional convolutional neural networks (3D CNNs) have demonstrated their outstanding classification accuracy for human action recognition (HAR). However, the large number of computations and parameters in 3D CNNs limits their deployability in real-life applications. To address this challenge, this paper adopts an algorithm-hardware co-design method by proposing an efficient 3D CNN building unit called 3D-1 bottleneck residual block (3D-1 BRB) at the algorithm level, and a corresponding FPGA-based hardware architecture called F-E3D at hardware level. Based on 3D-1 BRB, a novel 3D CNN model called E3DNet is developed, which achieves nearly 37 times reduction in model size and 5% improvement in accuracy compared to standard 3D CNNs on the UCF101 dataset. Together with several hardware optimizations, including 3D fused BRB, online blocking and kernel reuse, the proposed F-E3D is nearly 13 times faster than a previous FPGA design for 3D CNNs, with performance and accuracy comparable to other state-of-the-art 3D CNN models on GPU platforms while requiring only 7% of their energy consumption.

引用

页码：1 / 8

页数：8

共 50 条

[21] Human Action Recognition using 3D Convolutional Neural Networks with 3D Motion Cuboids in Surveillance Videos
Arunnehru, J.
Chamundeeswari, G.
Bharathi, S. Prasanna
INTERNATIONAL CONFERENCE ON ROBOTICS AND SMART MANUFACTURING (ROSMA2018), 2018, 133 : 471 - 477
[22] 3D CONVOLUTIONAL NEURAL NETWORK WITH MULTI-MODEL FRAMEWORK FOR ACTION RECOGNITION
Jing, Longlong
Ye, Yuancheng
Yang, Xiaodong
Tian, Yingli
2017 24TH IEEE INTERNATIONAL CONFERENCE ON IMAGE PROCESSING (ICIP), 2017, : 1837 - 1841
[23] Study of Human Motion Recognition Algorithm Based on Multichannel 3D Convolutional Neural Network
Ju, Yang
COMPLEXITY, 2021, 2021
[24] Enhanced 3D Action Recognition Based on Deep Neural Network
Park, Sungjoo
Kim, Dongchil
2022 THIRTEENTH INTERNATIONAL CONFERENCE ON UBIQUITOUS AND FUTURE NETWORKS (ICUFN), 2022, : 470 - 472
[25] Study on 3D Action Recognition Based on Deep Neural Network
Park, Sungjoo
Kim, Dongchil
2019 INTERNATIONAL CONFERENCE ON ELECTRONICS, INFORMATION, AND COMMUNICATION (ICEIC), 2019, : 309 - 311
[26] 3D convolutional neural network for object recognition: a review
Rahul Dev Singh
Ajay Mittal
Rajesh K. Bhatia
Multimedia Tools and Applications, 2019, 78 : 15951 - 15995
[27] A 3D Tensor Representation of Speech and 3D Convolutional Neural Network for Emotion Recognition
Mohammad Reza Falahzadeh
Fardad Farokhi
Ali Harimi
Reza Sabbaghi-Nadooshan
Circuits, Systems, and Signal Processing, 2023, 42 : 4271 - 4291
[28] A 3D Tensor Representation of Speech and 3D Convolutional Neural Network for Emotion Recognition
Falahzadeh, Mohammad Reza
Farokhi, Fardad
Harimi, Ali
Sabbaghi-Nadooshan, Reza
CIRCUITS SYSTEMS AND SIGNAL PROCESSING, 2023, 42 (07) : 4271 - 4291
[29] 3D convolutional neural network for object recognition: a review
Singh, Rahul Dev
Mittal, Ajay
Bhatia, Rajesh K.
MULTIMEDIA TOOLS AND APPLICATIONS, 2019, 78 (12) : 15951 - 15995
[30] 3D Human Motion Synthesis Based on Convolutional Neural Network
Zhou, Dongsheng
Feng, Xinzhu
Yi, Pengfei
Yang, Xin
Zhang, Qiang
Wei, Xiaopeng
Yang, Deyun
IEEE ACCESS, 2019, 7 : 66325 - 66335

← 1 2 3 4 5 →