This work presents an advanced pipeline to first separate audio and then give a summary of the conversation. The proposed model combines SepFormer, ConvTasNet, and adaptive noise reduction techniques to isolate speech from two-speaker mixed audio, reduce background noise, and amplify the primary speaker’s voice. This hybrid approach gives better results than each of the two models used on their own, without significant increase in computational cost. Once trained, the system delivers rapid, accurate audio separation and transcription. Performance evaluation is done using standard metrics, including Signal-to-Distortion Ratio (SDR), Signal-to-Interference Ratio (SIR), and Signal-to-Artefacts Ratio (SAR) and Scale-Invariant SNR (SI-SNR) and it demonstrates the effectiveness of the proposed model. The model yields an average SDR, SIR, SAR and SI-SNR of 24.6, 24.5, 24.5 and 21.9935 respectively which shows its capability in improving speech clarity while maintaining efficiency.

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

Multi-speaker Speech Processing in Noisy Environments: A Hybrid Model for Source Separation and Summarization

  • Satvik Raghav,
  • B. M. Vikhyath,
  • Raja Karthikeya,
  • S. Lalitha

摘要

This work presents an advanced pipeline to first separate audio and then give a summary of the conversation. The proposed model combines SepFormer, ConvTasNet, and adaptive noise reduction techniques to isolate speech from two-speaker mixed audio, reduce background noise, and amplify the primary speaker’s voice. This hybrid approach gives better results than each of the two models used on their own, without significant increase in computational cost. Once trained, the system delivers rapid, accurate audio separation and transcription. Performance evaluation is done using standard metrics, including Signal-to-Distortion Ratio (SDR), Signal-to-Interference Ratio (SIR), and Signal-to-Artefacts Ratio (SAR) and Scale-Invariant SNR (SI-SNR) and it demonstrates the effectiveness of the proposed model. The model yields an average SDR, SIR, SAR and SI-SNR of 24.6, 24.5, 24.5 and 21.9935 respectively which shows its capability in improving speech clarity while maintaining efficiency.