Active Speaker Detection (ASD) plays a vital role in scene understanding tasks, aiming to determine whether individuals in a scene are speaking. It has broad applications in areas such as speaker diarization and speaker tracking. Mainstream approaches typically rely on facial images and audio spectrograms to make predictions. In this paper, we explore the ASD task in fisheye meeting scenes and highlight the observation that participants tend to gaze at the active speaker. Based on this insight, we propose a gaze-driven multimodal ASD method. Specifically, we aggregate the gaze directions of all participants to construct a scene-level gaze field. A dedicated feature extraction branch then captures participants’ attention patterns, enhancing ASD accuracy. Furthermore, to address facial distortions caused by fisheye cameras, we employ Deformable Convolution networks (DConv). And we use Time Delay Neural Networks (TDNN) to extract temporal features from audio sequences. We evaluate our method on the FisheyeMeeting dataset, which targets multi-speaker detection in real-world meeting scenes. Experimental results demonstrate that our method significantly improves ASD accuracy in real-world meeting scenes.

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

Gaze-Driven Active Speaker Detection in Meetings

  • Weiwei Jiang,
  • Long Rao,
  • Gaole Dai,
  • Yifan Wu,
  • Wei Xu

摘要

Active Speaker Detection (ASD) plays a vital role in scene understanding tasks, aiming to determine whether individuals in a scene are speaking. It has broad applications in areas such as speaker diarization and speaker tracking. Mainstream approaches typically rely on facial images and audio spectrograms to make predictions. In this paper, we explore the ASD task in fisheye meeting scenes and highlight the observation that participants tend to gaze at the active speaker. Based on this insight, we propose a gaze-driven multimodal ASD method. Specifically, we aggregate the gaze directions of all participants to construct a scene-level gaze field. A dedicated feature extraction branch then captures participants’ attention patterns, enhancing ASD accuracy. Furthermore, to address facial distortions caused by fisheye cameras, we employ Deformable Convolution networks (DConv). And we use Time Delay Neural Networks (TDNN) to extract temporal features from audio sequences. We evaluate our method on the FisheyeMeeting dataset, which targets multi-speaker detection in real-world meeting scenes. Experimental results demonstrate that our method significantly improves ASD accuracy in real-world meeting scenes.