Accurate word boundary segmentation is a critical component for applications such as transcription, subtitling, and automatic speech recognition (ASR). In this paper, we present the Speech Recognition with Binary Search (SRwBS) algorithm, which integrates speech recognition with a binary search method to improve word boundary detection. Unlike traditional segmentation approaches, SRwBS refines word boundary estimations after speech recognition, offering improved segmentation precision. To evaluate SRwBS, we created a custom dataset with audio clips and word-level timing annotations. While our results show that SRwBS performs well in terms of precision, the recall remains relatively low, particularly in fast or complex speech, indicating that the effectiveness of the method is still heavily influenced by the quality of the underlying speech recognition system. Furthermore, although the processing time for segmenting audio clips was not real-time, we introduced the Faster SRwBS variant to optimize processing speed. While Faster SRwBS significantly improves the processing time, it still does not operate in real-time. This approach demonstrates significant potential for enhancing word boundary segmentation in ASR applications. Future work will focus on improving segmentation accuracy, real-time processing, and expanding the system’s capabilities across languages and domains.

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

Speech Recognition with Binary Search for Word Boundary Segmentation

  • Magzhan Kairanbay,
  • Reem Alhayki,
  • Fatema AlShaikh

摘要

Accurate word boundary segmentation is a critical component for applications such as transcription, subtitling, and automatic speech recognition (ASR). In this paper, we present the Speech Recognition with Binary Search (SRwBS) algorithm, which integrates speech recognition with a binary search method to improve word boundary detection. Unlike traditional segmentation approaches, SRwBS refines word boundary estimations after speech recognition, offering improved segmentation precision. To evaluate SRwBS, we created a custom dataset with audio clips and word-level timing annotations. While our results show that SRwBS performs well in terms of precision, the recall remains relatively low, particularly in fast or complex speech, indicating that the effectiveness of the method is still heavily influenced by the quality of the underlying speech recognition system. Furthermore, although the processing time for segmenting audio clips was not real-time, we introduced the Faster SRwBS variant to optimize processing speed. While Faster SRwBS significantly improves the processing time, it still does not operate in real-time. This approach demonstrates significant potential for enhancing word boundary segmentation in ASR applications. Future work will focus on improving segmentation accuracy, real-time processing, and expanding the system’s capabilities across languages and domains.