Extraction and Classification of Diving Clips from Continuous Video Footage

被引：16

作者：

Nibali, Aiden ^{[1
]}

He, Zhen ^{[1
]}

Morgan, Stuart ^{[1
,2
]}

Greenwood, Daniel ^{[2
]}

机构：

[1] La Trobe Univ, Bundoora, Vic, Australia

[2] Australian Inst Sport, Bruce, Australia

来源：

2017 IEEE CONFERENCE ON COMPUTER VISION AND PATTERN RECOGNITION WORKSHOPS (CVPRW) | 2017年

关键词：

ACTION RECOGNITION;

D O I：

10.1109/CVPRW.2017.18

中图分类号：

TP18 [人工智能理论];

学科分类号：

081104 ; 0812 ; 0835 ; 1405 ;

摘要：

Due to recent advances in technology, the recording and analysis of video data has become an increasingly common component of athlete training programmes. Today it is incredibly easy and affordable to set up a fixed camera and record athletes in a wide range of sports, such as diving, gymnastics, golf, tennis, etc. However, the manual analysis of the obtained footage is a time-consuming task which involves isolating actions of interest and categorizing them using domain-specific knowledge. In order to automate this kind of task, three challenging sub-problems are often encountered: 1) temporally cropping events/actions of interest from continuous video; 2) tracking the object of interest; and 3) classifying the events/actions of interest. Most previous work has focused on solving just one of the above sub-problems in isolation. In contrast, this paper provides a complete solution to the overall action monitoring task in the context of a challenging real-world exemplar. Specifically, we address the problem of diving classification. This is a challenging problem since the person (diver) of interest typically occupies fewer than 1% of the pixels in each frame. The model is required to learn the temporal boundaries of a dive, even though other divers and bystanders may be in view. Finally, the model must be sensitive to subtle changes in body pose over a large number of frames to determine the classification code. We provide effective solutions to each of the sub-problems which combine to provide a highly functional solution to the task as a whole. The techniques proposed can be easily generalized to video footage recorded from other sports.

引用

页码：94 / 104

页数：11

共 49 条

[1]

[Anonymous], 2015, P ADV NEURAL INFORM

[2]

[Anonymous], 2015, 1511 ARXIV

[3]

[Anonymous], 2015, ARXIV150206796

[4]

[Anonymous], 2015, ARXIV151106984

[5] Speeded-Up Robust Features (SURF) [J].

Bay, Herbert ;

Ess, Andreas ;

Tuytelaars, Tinne ;

Van Gool, Luc .

COMPUTER VISION AND IMAGE UNDERSTANDING, 2008, 110 (03) :346-359

[6]

Comaniciu D, 2000, PROC CVPR IEEE, P142, DOI 10.1109/CVPR.2000.854761

[7] Histograms of oriented gradients for human detection [J].

Dalal, N ;

Triggs, B .

2005 IEEE COMPUTER SOCIETY CONFERENCE ON COMPUTER VISION AND PATTERN RECOGNITION, VOL 1, PROCEEDINGS, 2005, :886-893

[8]

Donahue J, 2015, PROC CVPR IEEE, P2625, DOI 10.1109/CVPR.2015.7298878

[9] RANDOM SAMPLE CONSENSUS - A PARADIGM FOR MODEL-FITTING WITH APPLICATIONS TO IMAGE-ANALYSIS AND AUTOMATED CARTOGRAPHY [J].

FISCHLER, MA ;

BOLLES, RC .

COMMUNICATIONS OF THE ACM, 1981, 24 (06) :381-395

[10]

Girshick, 2015, P IEEE INT C COMP VI, DOI [10.1109/ICCV.2015.169, DOI 10.1109/ICCV.2015.169]

← 1 2 3 4 5 →