Dissecting neural computations in the human auditory pathway using deep neural networks for speech

被引：28

作者：

Li, Yuanning ^{[1
,4
,5
]}

Anumanchipalli, Gopala K. ^{[2
,3
]}

Mohamed, Abdelrahman

Chen, Peili ^{[4
,5
]}

Carney, Laurel H. ^{[6
]}

Lu, Junfeng ^{[7
,8
]}

Wu, Jinsong ^{[7
,8
]}

Chang, Edward F. ^{[1
,2
]}

机构：

[1] Univ Calif San Francisco, Dept Neurol Surg, San Francisco, CA 94115 USA

[2] Univ Calif San Francisco, Weill Inst Neurosci, San Francisco, CA 94143 USA

[3] Univ Calif Berkeley, Dept Elect Engn & Comp Sci, Berkeley, CA USA

[4] ShanghaiTech Univ, Sch Biomed Engn, Shanghai, Peoples R China

[5] Shanghai Tech Univ, State Key Lab Adv Med Mat & Devices, Shanghai, Peoples R China

[6] Univ Rochester, Dept Biomed Engn, Rochester, NY USA

[7] Fudan Univ, Huashan Hosp, Shanghai Med Coll, Neurol Surg Dept, Shanghai, Peoples R China

[8] Fudan Univ, Neurosurg Inst, Brain Funct Lab, Shanghai, Peoples R China

来源：

NATURE NEUROSCIENCE | 2023年 / 26卷 / 12期

基金：

中国国家自然科学基金;

关键词：

MODELS; CORTEX; REPRESENTATIONS; ORGANIZATION; CONNECTIONS; PERCEPTION; RESPONSES; MEANINGS; NEURONS; OBJECTS;

D O I：

10.1038/s41593-023-01468-4

中图分类号：

Q189 [神经科学];

学科分类号：

071006 ;

摘要：

The human auditory system extracts rich linguistic abstractions from speech signals. Traditional approaches to understanding this complex process have used linear feature-encoding models, with limited success. Artificial neural networks excel in speech recognition tasks and offer promising computational models of speech processing. We used speech representations in state-of-the-art deep neural network (DNN) models to investigate neural coding from the auditory nerve to the speech cortex. Representations in hierarchical layers of the DNN correlated well with the neural activity throughout the ascending auditory system. Unsupervised speech models performed at least as well as other purely supervised or fine-tuned models. Deeper DNN layers were better correlated with the neural activity in the higher-order auditory cortex, with computations aligned with phonemic and syllabic structures in speech. Accordingly, DNN models trained on either English or Mandarin predicted cortical responses in native speakers of each language. These results reveal convergence between DNN model representations and the biological auditory pathway, offering new approaches for modeling neural coding in the auditory cortex. Using direct intracranial recordings and modern speech AI models, Li and colleagues show representational and computational similarities between deep neural networks for self-supervised speech learning and the human auditory pathway.

引用

页码：2213 / 2225

页数：30

共 74 条

[1] Representations of Pitch and Timbre Variation in Human Auditory Cortex [J].