FRAMEWORK FOR EVALUATION OF SOUND EVENT DETECTION IN WEB VIDEOS

被引:0
|
作者
Badlani, Rohan [2 ]
Shah, Ankit [1 ]
Elizalde, Benjamin [1 ]
Kumar, Anurag [1 ]
Raj, Bhiksha [1 ]
机构
[1] Carnegie Mellon Univ, Language Technol Inst, Pittsburgh, PA 15213 USA
[2] BITS Pilani, Dept Comp Sci, Hyderabad, Telangana, India
来源
2018 IEEE INTERNATIONAL CONFERENCE ON ACOUSTICS, SPEECH AND SIGNAL PROCESSING (ICASSP) | 2018年
关键词
Sound Event Detection; Convolutional Neural Network; Large-Scale audio event detection; Video Content Analysis;
D O I
暂无
中图分类号
O42 [声学];
学科分类号
070206 ; 082403 ;
摘要
The largest source of sound events is web videos. Most videos lack sound event labels at segment level, however, a significant number of them do respond to text queries, from a match found using metadata by search engines. In this paper we explore the extent to which a search query can be used as the true label for detection of sound events in videos. We present a framework for large-scale sound event recognition on web videos. The framework crawls videos using search queries corresponding to 78 sound event labels drawn from three datasets. The datasets are used to train three classifiers, and we obtain a prediction on 3.7 million web video segments. We evaluated performance using the search query as true label and compare it with human labeling. Both types of ground truth exhibited close performance, to within 10%, and similar performance trend with increasing number of evaluated segments. Hence, our experiments show potential for using search query as a preliminary true label for sound event recognition in web videos.
引用
收藏
页码:3096 / 3100
页数:5
相关论文
共 50 条
  • [21] IMPACT OF SOUND DURATION AND INACTIVE FRAMES ON SOUND EVENT DETECTION PERFORMANCE
    Imoto, Keisuke
    Mishima, Sakiko
    Arai, Yumi
    Kondo, Reishi
    2021 IEEE INTERNATIONAL CONFERENCE ON ACOUSTICS, SPEECH AND SIGNAL PROCESSING (ICASSP 2021), 2021, : 860 - 864
  • [22] Sound Event Detection and Localization with Distance Estimation
    Krause, Daniel Aleksander
    Politis, Archontis
    Mesaros, Annamaria
    32ND EUROPEAN SIGNAL PROCESSING CONFERENCE, EUSIPCO 2024, 2024, : 286 - 290
  • [23] A Task-Specific Meta-Learning Framework for Few-Shot Sound Event Detection
    Zhang, Tianyang
    Yang, Liping
    Gu, Xiaohua
    Wang, Yuyang
    2022 IEEE 24TH INTERNATIONAL WORKSHOP ON MULTIMEDIA SIGNAL PROCESSING (MMSP), 2022,
  • [24] BEYOND THE DCASE 2017 CHALLENGE ON RARE SOUND EVENT DETECTION: A PROPOSAL FOR A MORE REALISTIC TRAINING AND TEST FRAMEWORK
    Baumann, Jan
    Lohrenz, Timo
    Roy, Alexander
    Fingscheidt, Tim
    2020 IEEE INTERNATIONAL CONFERENCE ON ACOUSTICS, SPEECH, AND SIGNAL PROCESSING, 2020, : 611 - 615
  • [25] Sound Event Detection Using Multiple Optimized Kernels
    Xia, Xianjun
    Tognerie, Roberto
    Sohel, Ferdous
    Zhaoe, Yuanjun
    Huang, Defeng
    IEEE-ACM TRANSACTIONS ON AUDIO SPEECH AND LANGUAGE PROCESSING, 2020, 28 (28) : 1745 - 1754
  • [26] Parallel Capsule Neural Networks for Sound Event Detection
    Liang, Kai-Wen
    Tseng, Yu-Hao
    Chang, Pao-Chi
    2019 ASIA-PACIFIC SIGNAL AND INFORMATION PROCESSING ASSOCIATION ANNUAL SUMMIT AND CONFERENCE (APSIPA ASC), 2019, : 1933 - 1936
  • [27] A survey of Deep Learning for Polyphonic Sound event detection
    Dang, An
    Vu, Toan H.
    Wang, Jia-Ching
    PROCEEDINGS OF THE 2017 INTERNATIONAL CONFERENCE ON ORANGE TECHNOLOGIES (ICOT), 2017, : 75 - 78
  • [28] Sound Event Detection with Depthwise Separable and Dilated Convolutions
    Drossos, Konstantinos
    Mimilakis, Stylianos, I
    Gharib, Shayan
    Li, Yanxiong
    Virtanen, Tuomas
    2020 INTERNATIONAL JOINT CONFERENCE ON NEURAL NETWORKS (IJCNN), 2020,
  • [29] SOUND EVENT DETECTION GUIDED BY SEMANTIC CONTEXTS OF SCENES
    Tonami, Noriyuki
    Imoto, Keisuke
    Nagase, Ryotaro
    Okamoto, Yuki
    Fukumori, Takahiro
    Yamashita, Yoichi
    2022 IEEE INTERNATIONAL CONFERENCE ON ACOUSTICS, SPEECH AND SIGNAL PROCESSING (ICASSP), 2022, : 801 - 805
  • [30] PEER COLLABORATIVE LEARNING FOR POLYPHONIC SOUND EVENT DETECTION
    Endo, Hayato
    Nishizaki, Hiromitsu
    2022 IEEE INTERNATIONAL CONFERENCE ON ACOUSTICS, SPEECH AND SIGNAL PROCESSING (ICASSP), 2022, : 826 - 830