The multisensory cocktail party problem in adults: The effects of the configural talking-face template and of facial, vocal, and linguistic identity cues on its solution
Abstract
Social communication often involves competing talkers and, thus, gives rise to the multisensory cocktail party problem (MCPP). To solve the MCPP, perceivers must bind, integrate, and segregate each talker’s auditory (A) and visual (V) speech streams. Audiovisual (AV) temporal synchrony is a known and powerful multisensory segregation cue and facial and vocal identity cues are known to facilitate segregation. Using eye tracking, we investigated whether perceptual segregation of multiple talking faces depends on the configural talking-face template (Experiment 1) and whether facial, vocal, and linguistic identity cues associated with different talkers can facilitate perceptual segregation (Experiment 2). Results indicated (a) that AV synchrony-based perceptual segregation of multiple talking faces does not depend on the configural talking-face template, (b) that linguistic identity cues facilitated synchrony-based segregation of upright talking faces independently of facial and vocal identity cues, (c) that perceptual segregation was driven primarily by AV speech processing as evidenced by greater selective attention to the talkers’ mouth, and (d) that, as evidenced by pupillary dilation, the type and number of AV congruence cues that had to be processed for successful segregation was associated with differential cognitive effort. Overall, these findings add to existing evidence indicating that adults’ solution of the MCPP relies on AV temporal synchrony cues and on facial, vocal, and linguistic identity congruence cues but not on the canonical configural talking-face template.
Article Details
Authors (2)
David J. Lewkowicz
Julia McClellan