<p>In this paper, we address the complex problem of detecting overlapping speech segments, a key challenge in speech processing with applications in speaker diarization, automatic transcription, and multi-speaker recognition systems. Traditional approaches, often relying on exponential loss functions within the AdaBoost framework, struggle to maintain robustness in noisy or imbalanced data environments. To enhance detection accuracy, we propose a novel AdaBoost variant—SLFAdaBoost—utilizing a squared loss function specifically tailored for overlapping speech detection. This model demonstrates superior stability and convergence, significantly outperforming conventional methods. Experimental evaluation on the NIST 2005 corpus revealed that SLFAdaBoost achieves an accuracy range of 92.5% to 98.6%, with F1 scores between 95.4% and 99.2%, notably surpassing standard AdaBoost and Random Forest classifiers. Additionally, our model shows resilience in noisy conditions, maintaining precision and recall metrics at higher noise levels than other classifiers. These results underscore SLFAdaBoost's capacity to handle intricate data patterns, offering a more robust solution for real-world speech processing applications where overlapping segments are prevalent. This contribution provides an efficient, high-performing model, advancing the capabilities of ensemble methods in complex audio environments.</p>

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

A novel approach to deriving adaboost classifier weights using squared loss function for overlapping speech detection

  • Nassim Asbai,
  • Hadjer Bounazou,
  • Sihem Zitouni

摘要

In this paper, we address the complex problem of detecting overlapping speech segments, a key challenge in speech processing with applications in speaker diarization, automatic transcription, and multi-speaker recognition systems. Traditional approaches, often relying on exponential loss functions within the AdaBoost framework, struggle to maintain robustness in noisy or imbalanced data environments. To enhance detection accuracy, we propose a novel AdaBoost variant—SLFAdaBoost—utilizing a squared loss function specifically tailored for overlapping speech detection. This model demonstrates superior stability and convergence, significantly outperforming conventional methods. Experimental evaluation on the NIST 2005 corpus revealed that SLFAdaBoost achieves an accuracy range of 92.5% to 98.6%, with F1 scores between 95.4% and 99.2%, notably surpassing standard AdaBoost and Random Forest classifiers. Additionally, our model shows resilience in noisy conditions, maintaining precision and recall metrics at higher noise levels than other classifiers. These results underscore SLFAdaBoost's capacity to handle intricate data patterns, offering a more robust solution for real-world speech processing applications where overlapping segments are prevalent. This contribution provides an efficient, high-performing model, advancing the capabilities of ensemble methods in complex audio environments.