FRAMEWORK FOR EVALUATION OF SOUND EVENT DETECTION IN WEB VIDEOS

被引:0
作者
Badlani, Rohan [2 ]
Shah, Ankit [1 ]
Elizalde, Benjamin [1 ]
Kumar, Anurag [1 ]
Raj, Bhiksha [1 ]
机构
[1] Carnegie Mellon Univ, Language Technol Inst, Pittsburgh, PA 15213 USA
[2] BITS Pilani, Dept Comp Sci, Hyderabad, Telangana, India
来源
2018 IEEE INTERNATIONAL CONFERENCE ON ACOUSTICS, SPEECH AND SIGNAL PROCESSING (ICASSP) | 2018年
关键词
Sound Event Detection; Convolutional Neural Network; Large-Scale audio event detection; Video Content Analysis;
D O I
暂无
中图分类号
O42 [声学];
学科分类号
070206 ; 082403 ;
摘要
The largest source of sound events is web videos. Most videos lack sound event labels at segment level, however, a significant number of them do respond to text queries, from a match found using metadata by search engines. In this paper we explore the extent to which a search query can be used as the true label for detection of sound events in videos. We present a framework for large-scale sound event recognition on web videos. The framework crawls videos using search queries corresponding to 78 sound event labels drawn from three datasets. The datasets are used to train three classifiers, and we obtain a prediction on 3.7 million web video segments. We evaluated performance using the search query as true label and compare it with human labeling. Both types of ground truth exhibited close performance, to within 10%, and similar performance trend with increasing number of evaluated segments. Hence, our experiments show potential for using search query as a preliminary true label for sound event recognition in web videos.
引用
收藏
页码:3096 / 3100
页数:5
相关论文
共 50 条
  • [41] A Multi-Task Learning Framework for Sound Event Detection using High-level Acoustic Characteristics of Sounds
    Khandelwal, Tanmay
    Das, Rohan Kumar
    INTERSPEECH 2023, 2023, : 1214 - 1218
  • [42] Master-Teacher-Student: A Weakly Labelled Semi-Supervised Framework for Audio Tagging and Sound Event Detection
    Liu, Yuzhuo
    Chen, Hangting
    Zhao, Qingwei
    Zhang, Pengyuan
    IEICE TRANSACTIONS ON INFORMATION AND SYSTEMS, 2022, E105D (04) : 828 - 831
  • [43] Sound Event Detection Utilizing Graph Laplacian Regularization with Event Co-Occurrence
    Imoto, Keisuke
    Kyochi, Seisuke
    IEICE TRANSACTIONS ON INFORMATION AND SYSTEMS, 2020, E103D (09) : 1971 - 1977
  • [44] Voice activity detection in the wild via weakly supervised sound event detection
    Chen, Yefei
    Dinkel, Heinrich
    Wu, Mengyue
    Yu, Kai
    INTERSPEECH 2020, 2020, : 3665 - 3669
  • [45] Active Few-Shot Learning for Sound Event Detection
    Wang, Yu
    Cartwright, Mark
    Bello, Juan Pablo
    INTERSPEECH 2022, 2022, : 1551 - 1555
  • [46] SELF-TRAINING FOR SOUND EVENT DETECTION IN AUDIO MIXTURES
    Park, Sangwook
    Bellur, Ashwin
    Han, David K.
    Elhilali, Mounya
    2021 IEEE INTERNATIONAL CONFERENCE ON ACOUSTICS, SPEECH AND SIGNAL PROCESSING (ICASSP 2021), 2021, : 341 - 345
  • [47] Sound Event Detection for Human Safety and Security in Noisy Environments
    Neri, Michael
    Battisti, Federica
    Neri, Alessandro
    Carli, Marco
    IEEE ACCESS, 2022, 10 : 134230 - 134240
  • [48] Enhance Temporal Relations in Audio Captioning with Sound Event Detection
    Xie, Zeyu
    Xu, Xuenan
    Wu, Mengyue
    Yu, Kai
    INTERSPEECH 2023, 2023, : 4179 - 4183
  • [49] Multi Model-Based Distillation for Sound Event Detection
    Fu, Yingwei
    Xu, Kele
    Mi, Haibo
    Kong, Qiuqiang
    Wang, Dezhi
    Wang, Huaimin
    Hong, Tie
    IEICE TRANSACTIONS ON INFORMATION AND SYSTEMS, 2019, E102D (10) : 2055 - 2058
  • [50] Weak Supervised Sound Event Detection Based on Puzzle CAM
    Cai, Xichang
    Gan, Yanggang
    Wu, Menglong
    Wu, Juan
    IEEE ACCESS, 2023, 11 : 89290 - 89297