Factorized visual representations in the primate visual system and deep neural networks

被引:0
作者
Lindsey, Jack W. [1 ,2 ]
Issa, Elias B. [1 ,2 ]
机构
[1] Columbia Univ, Zuckerman Mind Brain Behav Inst, New York, NY 10027 USA
[2] Columbia Univ, Dept Neurosci, New York, NY 10027 USA
关键词
visual cortex; deep neural networks; neurophysiology; visual scenes; object recognition; fMRI; OBJECT; SELECTIVITY; INCREASES; FRAMEWORK; GEOMETRY;
D O I
10.7554/eLife.91685; 10.7554/eLife.91685.3.sa1; 10.7554/eLife.91685.3.sa2; 10.7554/eLife.91685.3.sa3
中图分类号
Q [生物科学];
学科分类号
07 ; 0710 ; 09 ;
摘要
Object classification has been proposed as a principal objective of the primate ventral visual stream and has been used as an optimization target for deep neural network models (DNNs) of the visual system. However, visual brain areas represent many different types of information, and optimizing for classification of object identity alone does not constrain how other information may be encoded in visual representations. Information about different scene parameters may be discarded altogether ('invariance'), represented in non-interfering subspaces of population activity ('factorization') or encoded in an entangled fashion. In this work, we provide evidence that factorization is a normative principle of biological visual representations. In the monkey ventral visual hierarchy, we found that factorization of object pose and background information from object identity increased in higher-level regions and strongly contributed to improving object identity decoding performance. We then conducted a large-scale analysis of factorization of individual scene parameters - lighting, background, camera viewpoint, and object pose - in a diverse library of DNN models of the visual system. Models which best matched neural, fMRI, and behavioral data from both monkeys and humans across 12 datasets tended to be those which factorized scene parameters most strongly. Notably, invariance to these parameters was not as consistently associated with matches to neural and behavioral data, suggesting that maintaining non-class information in factorized activity subspaces is often preferred to dropping it altogether. Thus, we propose that factorization of visual scene information is a widely used strategy in brains and DNN models thereof.
引用
收藏
页数:20
相关论文
共 56 条
[31]   The ventral visual pathway: an expanded neural framework for the processing of object quality [J].
Kravitz, Dwight J. ;
Saleem, Kadharbatcha S. ;
Baker, Chris I. ;
Ungerleider, Leslie G. ;
Mishkin, Mortimer .
TRENDS IN COGNITIVE SCIENCES, 2013, 17 (01) :26-49
[32]   A new neural framework for visuospatial processing [J].
Kravitz, Dwight J. ;
Saleem, Kadharbatcha S. ;
Baker, Chris I. ;
Mishkin, Mortimer .
NATURE REVIEWS NEUROSCIENCE, 2011, 12 (04) :217-230
[33]   Representational geometry: integrating cognition, computation, and the brain [J].
Kriegeskorte, Nikolaus ;
Kievit, Rogier A. .
TRENDS IN COGNITIVE SCIENCES, 2013, 17 (08) :401-412
[34]   ImageNet Classification with Deep Convolutional Neural Networks [J].
Krizhevsky, Alex ;
Sutskever, Ilya ;
Hinton, Geoffrey E. .
COMMUNICATIONS OF THE ACM, 2017, 60 (06) :84-90
[35]   Parallel, multi-stage processing of colors, faces and shapes in macaque inferior temporal cortex [J].
Lafer-Sousa, Rosa ;
Conway, Bevil R. .
NATURE NEUROSCIENCE, 2013, 16 (12) :1870-1878
[36]  
Lee YJ, 2012, PROC CVPR IEEE, P1346, DOI 10.1109/CVPR.2012.6247820
[37]   Norm-based face encoding by single neurons in the monkey inferotemporal cortex [J].
Leopold, David A. ;
Bondar, Igor V. ;
Giese, Martin A. .
NATURE, 2006, 442 (7102) :572-575
[38]  
Linsley D, 2023, Arxiv, DOI arXiv:2306.03779
[39]   Optimal Degrees of Synaptic Connectivity [J].
Litwin-Kumar, Ashok ;
Harris, Kameron Decker ;
Axel, Richard ;
Sompolinsky, Haim ;
Abbott, L. F. .
NEURON, 2017, 93 (05) :1153-+
[40]   Simple Learned Weighted Sums of Inferior Temporal Neuronal Firing Rates Accurately Predict Human Core Object Recognition Performance [J].
Majaj, Najib J. ;
Hong, Ha ;
Solomon, Ethan A. ;
DiCarlo, James J. .
JOURNAL OF NEUROSCIENCE, 2015, 35 (39) :13402-13418