<p>The crucial role of signal processing in various machine learning-based speech applications has been amplified by advancements in deep learning techniques. However, accurately detecting voice activity in noisy conditions remains a major challenge, particularly when using high-dimensional spectro-temporal representations that require substantial processing power and learning capacity. This study aims to overcome this challenge by proposing a novel deep learning-based voice activity detection (VAD) method that integrates convolutional and recurrent architectures. Specifically, we introduce a hybrid ResNet50 and long short-term memory (LSTM) network within a spectro-temporal domain. The spectro-temporal domain offers a detailed visualization and analysis tool by capturing combined frequency and temporal features. However, its high dimensionality poses significant processing and learning challenges. To mitigate this, we introduce a sparsity-based dimension reduction algorithm that compresses the data while retaining essential VAD-relevant information. Our approach effectively distinguishes speech and non-speech segments using time-frequency features and a hybrid flexible classifier. It adapts to diverse noise types, including stationary, non-stationary, and periodic noises, and provides flexibility in selecting appropriate deep learning models for specific applications. Extensive comparisons with baseline and state-of-the-art methods demonstrate substantial improvements in speech enhancement, particularly in noisy conditions.</p>

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

A ResNet-LSTM hybrid deep learning approach for voice activity detection in the spectro-temporal domain

  • Samira Mavaddati,
  • Mohammad Razavi

摘要

The crucial role of signal processing in various machine learning-based speech applications has been amplified by advancements in deep learning techniques. However, accurately detecting voice activity in noisy conditions remains a major challenge, particularly when using high-dimensional spectro-temporal representations that require substantial processing power and learning capacity. This study aims to overcome this challenge by proposing a novel deep learning-based voice activity detection (VAD) method that integrates convolutional and recurrent architectures. Specifically, we introduce a hybrid ResNet50 and long short-term memory (LSTM) network within a spectro-temporal domain. The spectro-temporal domain offers a detailed visualization and analysis tool by capturing combined frequency and temporal features. However, its high dimensionality poses significant processing and learning challenges. To mitigate this, we introduce a sparsity-based dimension reduction algorithm that compresses the data while retaining essential VAD-relevant information. Our approach effectively distinguishes speech and non-speech segments using time-frequency features and a hybrid flexible classifier. It adapts to diverse noise types, including stationary, non-stationary, and periodic noises, and provides flexibility in selecting appropriate deep learning models for specific applications. Extensive comparisons with baseline and state-of-the-art methods demonstrate substantial improvements in speech enhancement, particularly in noisy conditions.