English Speech Segmentation and Keyword Extraction Based on Deep Learning
摘要
This article suggests a deep learning-based English speech segmentation and keyword extraction technique to enhance the precision and effectiveness of speech processing. The study utilizes a blend of CNN, BiLSTM, and self-attention mechanism for successful speech signal segmentation, keying, and word recognition through data preparation, feature extraction, model training, and optimization techniques. Experimental results indicate that the suggested model performs well in terms of accuracy and robustness for both speech segmentation and keyword extraction tasks, outperforming traditional methods by a significant margin. This study also enhances the model’s ability to generalize by employing various data augmentation methods, demonstrating its efficacy across various background noise scenarios. By comparing with other techniques, the proposed model is confirmed to be superior in the area of speech processing, which offers guidance and assistance for future intelligent speech recognition systems.