On the Use of Cross-Attentive Fusion Techniques for Audio-Visual Speaker Verification
摘要
Audio-Visual (A-V) speaker verification has recently been gaining a lot of attention owing to the closely associated audio and visual cues. Though existing approaches based on A-V fusion showed improvement over unimodal systems, its potential for speaker verification is not fully exploited. In this contribution, we investigate the prospect of effectively capturing the synergic relationships across audio and visual modalities, which can play a vital role to significantly boost the performance of multimodal fusion over unimodal systems. More specifically, we present a comparative study of three variants of Cross-Attentive frameworks, namely Joint Cross-Attention (JCA), Recursive JCA (RJCA) and Cross-Modal Transformer (CMT), for multimodal fusion that can efficiently capture the intra-modal and/or inter-modal associations for A-V speaker verification. We carry out extensive experiments on the Voxceleb1 dataset to rigorously evaluate and compare the cross-attention-based models. Results indicate that effectively capturing intra- and/or inter-modal relationships across audio and visual modalities can significantly improve the performance of the A-V speaker verification system.