Speaker change detection (SCD) is the task of finding points in an audio stream where one speaker changes to another. Traditional approaches, such as Bayesian information criterion with Gaussian mixture models (BIC-GMM) and Kullback-Leibler with Gaussian mixture models (KL-GMM), use traditional audio features and are very sensitive to threshold values, which makes them hard to generalize for different datasets and conditions. In this paper, we develop a SCD approach that uses a large pre-trained audio model, which serves as a feature extractor directly from raw audio. By leveraging the knowledge encoded in the large audio model, this method captures rich and robust audio representations. The extracted features are then passed to a neural network layer to identify the change points. By using the audio representations, the proposed method overcomes the limitations of traditional techniques. We tested the method on the meeting audio corpus with various window segments. The results showed that the proposed method reduces both false alarm rate and miss detection rate in a balanced way. Compared to traditional methods, the proposed approach reduces false alarm and miss detection by more than 20% and achieves a 10–15% improvement over baseline models that do not use pre-trained audio models. These results showed that using pre-trained audio models provides a more reliable and balanced solution for speaker change detection.

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

Speaker Change Detection with Pre-trained Large Audio Model

  • Alymzhan Toleu,
  • Gulmira Tolegen,
  • Alexander Krassovitskiy,
  • Rustam Mussabayev,
  • Bagashar Zhumazhanov

摘要

Speaker change detection (SCD) is the task of finding points in an audio stream where one speaker changes to another. Traditional approaches, such as Bayesian information criterion with Gaussian mixture models (BIC-GMM) and Kullback-Leibler with Gaussian mixture models (KL-GMM), use traditional audio features and are very sensitive to threshold values, which makes them hard to generalize for different datasets and conditions. In this paper, we develop a SCD approach that uses a large pre-trained audio model, which serves as a feature extractor directly from raw audio. By leveraging the knowledge encoded in the large audio model, this method captures rich and robust audio representations. The extracted features are then passed to a neural network layer to identify the change points. By using the audio representations, the proposed method overcomes the limitations of traditional techniques. We tested the method on the meeting audio corpus with various window segments. The results showed that the proposed method reduces both false alarm rate and miss detection rate in a balanced way. Compared to traditional methods, the proposed approach reduces false alarm and miss detection by more than 20% and achieves a 10–15% improvement over baseline models that do not use pre-trained audio models. These results showed that using pre-trained audio models provides a more reliable and balanced solution for speaker change detection.