Double Compressed Wideband AMR Speech Detection Using Deep Neural Networks

被引：0

作者：

Buker, Aykut ^{[1
]}

Hanilci, Cemal ^{[1
]}

机构：

[1] Bursa Tech Univ, Dept Elect & Elect Engn, TR-16310 Bursa, Turkiye

来源：

CIRCUITS SYSTEMS AND SIGNAL PROCESSING | 2024年 / 43卷 / 7期

关键词：

Audio forensics; Wideband AMR codec; Double compressed AMR detection; Deep neural networks;

D O I：

10.1007/s00034-024-02668-4

中图分类号：

TM [电工技术]; TN [电子技术、通信技术];

学科分类号：

0808 ; 0809 ;

摘要：

Detecting double compressed (DC) speech signals is an important audio forensics task since it is highly related to the integrity and the authenticity of the recording. Adaptive multi-rate (AMR) speech codec is a popular audio compression technique specifically optimized for speech signals and it is a standard audio recording format in the vast majority of the smart phones. All of the previous studies addressing the detection of DC AMR signals report their findings for the speech signals compressed using the narrowband AMR codec (AMR-NB). Meanwhile, wideband AMR codec (AMR-WB) has been used by several mobile phone manufacturers, but DC AMR-WB speech signal detection performance remains unknown. To the best of our knowledge, this is the first study focusing on detecting the DC signals compressed using the AMR-WB speech codec. To this end, we propose three different deep neural network-based DC AMR-WB signal detection systems where the spectrogram representations of the speech signals are used as the input features. Experimental results conducted on TIMIT database provide several important findings regarding the DC AMR-WB speech detection. Firstly, DC AMR-WB detection is found to be a more challenging task than detecting the AMR-NB signals. For example, convolutional neural network (CNN)-based system yields 74.83% and 99.93% detection rates on AMR-WB and AMR-NB coded signals, respectively. Secondly, capturing the temporal information using long short-term memory (LSTM) network with the DC AMR-WB signal detection accuracy of 86.25% is found to be superior to the CNN system. Thirdly, combining the deep feature representations learned by CNN and LSTM networks further improves the performance. Fourthly, the detection rates are found to deteriorate when the signals are first encoded using different audio codecs prior to AMR-WB compression. Finally, applying score level or decision level fusion to the proposed three systems improves the detection rates, in general.

引用

页码：4528 / 4546

页数：19

共 50 条

[1] Deep convolutional neural networks for double compressed AMR audio detection
Buker, Aykut
Hanilci, Cemal
IET SIGNAL PROCESSING, 2021, 15 (04) : 265 - 280
[2] Double Compressed AMR Audio Detection Using Long-Term Features and Deep Neural Networks
Buker, Aykut
Hanilci, Cemal
2019 11TH INTERNATIONAL CONFERENCE ON ELECTRICAL AND ELECTRONICS ENGINEERING (ELECO 2019), 2019, : 590 - 594
[3] Detection of AMR double compression using compressed -domain speech features
Sampaio, Jose F. P.
Nascimento, Francisco A. de O.
FORENSIC SCIENCE INTERNATIONAL-DIGITAL INVESTIGATION, 2020, 33
[4] Exploring the Effectiveness of the Phase Features on Double Compressed AMR Speech Detection
Buker, Aykut
Hanilci, Cemal
APPLIED SCIENCES-BASEL, 2024, 14 (11):
[5] DETECTING DOUBLE COMPRESSED AMR AUDIO USING DEEP LEARNING
Luo, Da
Yang, Rui
Huang, Jiwu
2014 IEEE INTERNATIONAL CONFERENCE ON ACOUSTICS, SPEECH AND SIGNAL PROCESSING (ICASSP), 2014,
[6] Speech Activity Detection Using Deep Neural Networks
Shahsavari, Sajad
Sameti, Hossein
Hadian, Hossein
2017 25TH IRANIAN CONFERENCE ON ELECTRICAL ENGINEERING (ICEE), 2017, : 1564 - 1568
[7] Detection of Double Compressed AMR Audio Using Stacked Autoencoder
Luo, Da
Yang, Rui
Li, Bin
Huang, Jiwu
IEEE TRANSACTIONS ON INFORMATION FORENSICS AND SECURITY, 2017, 12 (02) : 432 - 444
[8] Speech Activity Detection on YouTube Using Deep Neural Networks
Ryant, Neville
Liberman, Mark
Yuan, Jiahong
14TH ANNUAL CONFERENCE OF THE INTERNATIONAL SPEECH COMMUNICATION ASSOCIATION (INTERSPEECH 2013), VOLS 1-5, 2013, : 728 - 731
[9] Enhanced speech emotion detection using deep neural networks
S. Lalitha
Shikha Tripathi
Deepa Gupta
International Journal of Speech Technology, 2019, 22 : 497 - 510
[10] Enhanced speech emotion detection using deep neural networks
Lalitha, S.
Tripathi, Shikha
Gupta, Deepa
INTERNATIONAL JOURNAL OF SPEECH TECHNOLOGY, 2019, 22 (03) : 497 - 510

← 1 2 3 4 5 →