Audio-Visual (A-V) speaker verification has recently been gaining a lot of attention owing to the closely associated audio and visual cues. Though existing approaches based on A-V fusion showed improvement over unimodal systems, its potential for speaker verification is not fully exploited. In this contribution, we investigate the prospect of effectively capturing the synergic relationships across audio and visual modalities, which can play a vital role to significantly boost the performance of multimodal fusion over unimodal systems. More specifically, we present a comparative study of three variants of Cross-Attentive frameworks, namely Joint Cross-Attention (JCA), Recursive JCA (RJCA) and Cross-Modal Transformer (CMT), for multimodal fusion that can efficiently capture the intra-modal and/or inter-modal associations for A-V speaker verification. We carry out extensive experiments on the Voxceleb1 dataset to rigorously evaluate and compare the cross-attention-based models. Results indicate that effectively capturing intra- and/or inter-modal relationships across audio and visual modalities can significantly improve the performance of the A-V speaker verification system.

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

On the Use of Cross-Attentive Fusion Techniques for Audio-Visual Speaker Verification

  • Jahangir Alam,
  • R. Gnana Praveen

摘要

Audio-Visual (A-V) speaker verification has recently been gaining a lot of attention owing to the closely associated audio and visual cues. Though existing approaches based on A-V fusion showed improvement over unimodal systems, its potential for speaker verification is not fully exploited. In this contribution, we investigate the prospect of effectively capturing the synergic relationships across audio and visual modalities, which can play a vital role to significantly boost the performance of multimodal fusion over unimodal systems. More specifically, we present a comparative study of three variants of Cross-Attentive frameworks, namely Joint Cross-Attention (JCA), Recursive JCA (RJCA) and Cross-Modal Transformer (CMT), for multimodal fusion that can efficiently capture the intra-modal and/or inter-modal associations for A-V speaker verification. We carry out extensive experiments on the Voxceleb1 dataset to rigorously evaluate and compare the cross-attention-based models. Results indicate that effectively capturing intra- and/or inter-modal relationships across audio and visual modalities can significantly improve the performance of the A-V speaker verification system.