A Query-by-Singing System for Retrieving Karaoke Music

被引:26
作者
Yu, Hung-Ming [1 ]
Tsai, Wei-Ho [2 ,3 ]
Wang, Hsin-Min
机构
[1] Acad Sinica, Inst Informat Sci, Taipei, Taiwan
[2] Natl Taipei Univ Technol, Dept Elect Engn, Taipei, Taiwan
[3] Natl Taipei Univ Technol, Grad Inst Comp & Commun Engn, Taipei, Taiwan
关键词
Bayesian information criterion; dynamic time warping; karaoke; music information retrieval; query-by-singing;
D O I
10.1109/TMM.2008.2007345
中图分类号
TP [自动化技术、计算机技术];
学科分类号
0812 ;
摘要
This paper investigates the problem of retrieving karaoke music using query-by-singing techniques. Unlike regular CD music, where the stereo sound involves two audio channels that usually sound the same, karaoke music encompasses two distinct channels in each track: one is a mixture of the lead vocals and background accompaniment, and the other consists of accompaniment only. Although the two audio channels are distinct, the accompaniments in the two channels often resemble each other. We exploit this characteristic to: i) infer the background accompaniment for the lead vocals from the accompaniment-only channel, so that the main melody underlying the lead vocals can be extracted more effectively; and ii) detect phrase onsets based on the Bayesian information criterion (BIC) to predict the onset points of a song where a user's sung query may begin, so that the similarity between the melodies of the query and the song can be examined more efficiently. To further refine extraction of the main melody, we propose correcting potential errors in the estimated sung notes by exploiting a composition characteristic of popular songs whereby the sung notes within a verse or chorus section usually vary no more than two octaves. In addition, to facilitate an efficient and accurate search of a large music database, we employ multiple-pass dynamic time warping (DTW) combined with multiple-level data abstraction (MLDA) to compare the similarities of melodies. The results of experiments conducted on a karaoke database comprised of 1071 popular songs demonstrate the feasibility of query-by-singing retrieval for karaoke music.
引用
收藏
页码:1626 / 1637
页数:12
相关论文
共 37 条
  • [1] [Anonymous], 2004, KDD WORKSHOP MINING
  • [2] [Anonymous], P INT C MUS INF RETR
  • [3] DANNENBERG RB, 2003, P INT C MUS INF RETR, P41
  • [4] A comparative evaluation of search techniques for query-by-humming using the MUSART testbed
    Dannenberg, Roger B.
    Birmingham, William P.
    Pardo, Bryan
    Hu, Ning
    Meek, Colin
    Tzanetakis, George
    [J]. JOURNAL OF THE AMERICAN SOCIETY FOR INFORMATION SCIENCE AND TECHNOLOGY, 2007, 58 (05): : 687 - 701
  • [5] Factors affecting music retrieval in query-by-melody
    De Mulder, Tom
    Martens, Jean-Pierre
    Pauws, Steffen
    Vignoli, Fabio
    Lesaffre, Micheline
    Leman, Marc
    De Baets, Bernard
    De Meyer, Hans
    [J]. IEEE TRANSACTIONS ON MULTIMEDIA, 2006, 8 (04) : 728 - 739
  • [6] Robust polyphonic music retrieval with N-grams
    Doraisamy, S
    Rüeger, S
    [J]. JOURNAL OF INTELLIGENT INFORMATION SYSTEMS, 2003, 21 (01) : 53 - 70
  • [7] DORAISAMY S, 2001, P INT S MUS INF RETR
  • [8] Duda A., 2007, P INT C MUS INF RETR
  • [9] EGGINK J, 2004, P INT C MUS INF RETR
  • [10] FENG Y, 2003, P ACM C RES DEV INF