<p>Our visual brain transforms small differences between images in the two eyes (binocular disparity) into coherent depth. Initially, neurons in the primary visual cortex (V1) compute the degrees of overlap between the left and right images to encode disparity. Such cross-correlation-like neurons respond to both binocularly matched and mismatched features. This ambiguous representation is refined along the visual pathway through a cross-matching computation involving additional nonlinear processing to filter out mismatches. How these representations are organized in the human visual cortex remains unclear. Using functional magnetic resonance imaging (fMRI), we show that areas V1–V3 exhibit stronger cross-correlation components, while V3A/B, V7, hV4, and hMT+ are inclined towards cross-matching. A deep neural network (DNN) trained for stereo vision undergoes a similar transformation across its layers, progressing through distinct phases that exploit dissimilar features to achieve coherent depth. This brain-DNN alignment demonstrates that human and artificial visual systems share a computational principle for robust 3D vision.</p>

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

Human and artificial visual systems share a computational principle for transforming binocular disparity into depth representation

  • Bayu Gautama Wundari,
  • Ichiro Fujita,
  • Hiroshi Ban

摘要

Our visual brain transforms small differences between images in the two eyes (binocular disparity) into coherent depth. Initially, neurons in the primary visual cortex (V1) compute the degrees of overlap between the left and right images to encode disparity. Such cross-correlation-like neurons respond to both binocularly matched and mismatched features. This ambiguous representation is refined along the visual pathway through a cross-matching computation involving additional nonlinear processing to filter out mismatches. How these representations are organized in the human visual cortex remains unclear. Using functional magnetic resonance imaging (fMRI), we show that areas V1–V3 exhibit stronger cross-correlation components, while V3A/B, V7, hV4, and hMT+ are inclined towards cross-matching. A deep neural network (DNN) trained for stereo vision undergoes a similar transformation across its layers, progressing through distinct phases that exploit dissimilar features to achieve coherent depth. This brain-DNN alignment demonstrates that human and artificial visual systems share a computational principle for robust 3D vision.