Variations in unisensory speech perception explain interindividual differences in McGurk illusion susceptibility
摘要
Face-to-face communication relies on integrating acoustic speech signals with corresponding facial articulations. Audiovisual integration abilities or deficits in typical and atypical populations are often assessed through their susceptibility to the McGurk illusion (i.e., their McGurk illusion rates). According to theories of normative Bayesian causal inference, observers integrate a visual /ga/ viseme and an auditory /ba/ phoneme weighted by their relative phonemic reliabilities into an illusory “da” percept. Consequently, McGurk illusion rates should be strongly influenced by observers’ categorical perception of the corresponding facial articulatory movements and the acoustic signals. Across three experiments we investigated the extent to which variability in the McGurk illusion rate across participants or stimuli (i.e., speakers) can be explained by the corresponding variations in the categorical perception of the unisensory auditory and visual components. Additionally, we investigated whether the McGurk illusion susceptibility is a stable trait across different testing sessions (i.e., days) and tasks. Consistent with the principles of Bayesian Causal Inference, our results demonstrate that observers’ tendency to (mis)perceive the auditory /ba/ and the visual /ga/ stimuli as “da” in unisensory contexts strongly predicts their McGurk illusion rates across both speakers and participants. Likewise, the stability in the McGurk illusion across sessions and tasks arises closely aligned with the corresponding stability of the unisensory auditory and visual categorical perception. Collectively, these findings highlight the importance of accounting for variations in unisensory performance and variability of materials (e.g., speakers) when using audiovisual illusions to assess audiovisual integration capability.