ARTSTREAM: a neural network model of auditory scene analysis and source segregation

被引:51
作者
Grossberg, S
Govindarajan, KK
Wyse, LL
Cohen, MA
机构
[1] Boston Univ, Dept Cognit & Neural Syst, Ctr Adapt Syst, Boston, MA 02215 USA
[2] SpeechWorks Int, Boston, MA 02111 USA
[3] Informat Technol Lab, Singapore 119613, Singapore
基金
美国国家科学基金会;
关键词
auditory scene analysis; streaming; cocktail party problem; pitch perception; spatial localization; neural network; resonance; adaptive resonance theory; spectral-pitch resonance;
D O I
10.1016/j.neunet.2003.10.002
中图分类号
TP18 [人工智能理论];
学科分类号
081104 ; 0812 ; 0835 ; 1405 ;
摘要
Multiple sound sources often contain harmonics that overlap and may be degraded by environmental noise. The auditory system is capable of teasing apart these sources into distinct mental objects, or streams. Such an 'auditory scene analysis' enables the brain to solve the cocktail party problem. A neural network model of auditory scene analysis, called the ARTSTREAM model, is presented to propose how the brain accomplishes this feat. The model clarifies how the frequency components that correspond to a given acoustic source may be coherently grouped to-ether into a distinct stream based on pitch and spatial location Cues. The model also clarifies how multiple streams may be distinguished and separated by the brain. Streams are formed as spectral-pitch resonances that emerge through feedback interactions between frequency-specific spectral representations of a sound Source and its pitch. First, the model transforms a sound into a spatial pattern of frequency-specific activation across a spectral stream layer. The sound has multiple parallel representations at this layer. A sound's spectral representation activates it bottom-up filter that is sensitive to the harmonics of the sound's pitch. This filter activates a pitch category which, in turn, activates a top-down expectation that is also sensitive to the harmonics of the pitch. Resonance develops when the spectral and pitch representations mutually reinforce one another. Resonance provides the coherence that allows one voice or instrument to be tracked through a noisy multiple source environment. Spectral components are suppressed if they do not match harmonics of the top-down expectation that is read-out by the selected pitch, thereby allowing another stream to capture these components, as in the 'old-plus-new heuristic' of Bregman. Multiple simultaneously occurring spectral-pitch resonances can hereby emerge. These resonance and matching mechanisms are specialized versions of Adaptive Resonance Theory, or ART, which clarifies how pitch representations can self-organize during learning of harmonic bottom-up filters and top-down expectations. The model also clarifies how spatial location cues can help to disambiguate two sources with similar spectral cues. Data are simulated front psychophysical grouping experiments, such as how a tone sweeping upwards in frequency creates a bounce percept by grouping with a downward sweeping tone due to proximity in frequency, even if noise replaces the tones at their intersection point. Illusory auditory percepts are also simulated, such as the auditory continuity illusion of a tone continuing through a noise burst even if the tone is not present during the noise, and the scale illusion of Deutsch whereby downward and upward scales presented alternately to the two ears are regrouped based on frequency proximity, leading to a bounce percept. Since related sorts of resonances have been used to quantitatively simulate psychophysical data about speech perception, the model strengthens the hypothesis that ART-like mechanisms are used at multiple levels of the auditory system. Proposals for developing the model to explain more complex streaming data are also provided. (C) 2004 Elsevier Ltd. All rights reserved.
引用
收藏
页码:511 / 536
页数:26
相关论文
共 81 条
[11]   ON THE FUSION OF SOUNDS REACHING DIFFERENT SENSE ORGANS [J].
BROADBENT, DE ;
LADEFOGED, P .
JOURNAL OF THE ACOUSTICAL SOCIETY OF AMERICA, 1957, 29 (06) :708-710
[12]   INTONATION AND THE PERCEPTUAL SEPARATION OF SIMULTANEOUS VOICES [J].
BROKX, JPL ;
NOOTEBOOM, SG .
JOURNAL OF PHONETICS, 1982, 10 (01) :23-36
[13]  
BROWN GJ, 1992, THESIS U SHEFFIELD
[14]   DISCRIMINATING BETWEEN COHERENT AND INCOHERENT FREQUENCY-MODULATION OF COMPLEX TONES [J].
CARLYON, RP .
JOURNAL OF THE ACOUSTICAL SOCIETY OF AMERICA, 1991, 89 (01) :329-340
[15]  
CARLYON RP, 1992, PROCESSING COMPLEX S
[16]  
Carpenter G., 1991, Pattern recognition by self-organizing neural networks
[17]   NORMAL AND AMNESIC LEARNING, RECOGNITION AND MEMORY BY A NEURAL MODEL OF CORTICO-HIPPOCAMPAL INTERACTIONS [J].
CARPENTER, GA ;
GROSSBERG, S .
TRENDS IN NEUROSCIENCES, 1993, 16 (04) :131-137
[18]   A MASSIVELY PARALLEL ARCHITECTURE FOR A SELF-ORGANIZING NEURAL PATTERN-RECOGNITION MACHINE [J].
CARPENTER, GA ;
GROSSBERG, S .
COMPUTER VISION GRAPHICS AND IMAGE PROCESSING, 1987, 37 (01) :54-115
[19]   THE PERCEPTUAL SEGREGATION OF SIMULTANEOUS AUDITORY SIGNALS - PULSE TRAIN SEGREGATION AND VOWEL SEGREGATION [J].
CHALIKIA, MH ;
BREGMAN, AS .
PERCEPTION & PSYCHOPHYSICS, 1989, 46 (05) :487-496
[20]   Neural dynamics of motion grouping: from aperture ambiguity to object speed and direction [J].
Chey, J ;
Grossberg, S ;
Mingolla, E .
JOURNAL OF THE OPTICAL SOCIETY OF AMERICA A-OPTICS IMAGE SCIENCE AND VISION, 1997, 14 (10) :2570-2594