Tagging Webcast Text in Baseball Videos by Video Segmentation and Text Alignment

被引:17
作者
Chiu, Chih-Yi [1 ]
Lin, Po-Chih [1 ]
Li, Sheng-Yang [1 ]
Tsai, Tsung-Han [1 ]
Tsai, Yu-Lung [1 ]
机构
[1] Natl Chiayi Univ, Dept Comp Sci & Informat Engn, Chiayi 60004, Taiwan
关键词
Event detection; genetic algorithm; multimodal fusion; unsupervised learning; video annotation; EVENT DETECTION; INFORMATION; FUSION;
D O I
10.1109/TCSVT.2012.2189478
中图分类号
TM [电工技术]; TN [电子技术、通信技术];
学科分类号
0808 ; 0809 ;
摘要
Sports video annotation, an active research area in the field of multimedia content understanding, is an essential process in applications, such as summarization, highlight extraction, event detection, and retrieval. This paper considers the issue in relation to the annotation of baseball videos. Conventional baseball video annotation frameworks are based primarily on video content analysis, such as scoreboard recognition and machine learning techniques, which require a substantial amount of human input to collect and organize training data. The performance of such frameworks might become unstable if they encounter audiovisual patterns not included in the training data. To address the issue, we propose a novel framework for baseball video annotation that aligns high-level webcast text with low-level video content. Several cues, which are derived from the video content and webcast text, are utilized for alignment by leveraging hierarchical agglomerative clustering and genetic algorithm optimization. In addition, we develop an unsupervised method to learn the pitch segment properties of baseball videos by Markov random walk, and thereby reduce the need for human intervention substantially. Our experiments demonstrate that the proposed framework yields a robust result against a variety of video content and enhances the automaticity in baseball video annotation.
引用
收藏
页码:999 / 1013
页数:15
相关论文
共 34 条
[1]  
Ando R., 2007, P 6 ACM INT C IM VID, P186
[2]  
[Anonymous], 2010, P 18 ACM INT C MULTI
[3]  
[Anonymous], P IEEE INT C MULT EX
[4]  
Ariki Y., 2003, P EUROSPEECH, P1453
[5]   Personalized abstraction of broadcasted American football video by highlight selection [J].
Babaguchi, N ;
Kawai, Y ;
Ogura, T ;
Kitahashi, T .
IEEE TRANSACTIONS ON MULTIMEDIA, 2004, 6 (04) :575-586
[6]  
Chang P, 2002, IEEE IMAGE PROC, P609
[7]  
Chao Liang, 2011, 2011 IEEE Conference on Computer Vision and Pattern Recognition (CVPR), P3377, DOI 10.1109/CVPR.2011.5995681
[8]   Fusion of audio and motion information on HMM-based highlight extraction for baseball games [J].
Cheng, Chih-Chieh ;
Hsu, Chiou-Ting .
IEEE TRANSACTIONS ON MULTIMEDIA, 2006, 8 (03) :585-599
[9]   Explicit semantic events detection and development of realistic applications for broadcasting baseball videos [J].
Chu, Wei-Ta ;
Wu, Ja-Ling .
MULTIMEDIA TOOLS AND APPLICATIONS, 2008, 38 (01) :27-50
[10]  
Cormen T.H., 2009, INTRO ALGORITHM