Advancements in deep neural networks have allowed automatic speech recognition (ASR) systems to attain human parity on several publicly available clean speech datasets. However, even state-of-the-art ASR systems experience performance degradation when confronted with adverse conditions, as a well-trained acoustic model is sensitive to variations in the speech domain, e.g., background noise. On the other hand, linguistic knowledge in large language models (LLMs) can be used to infer the meaning of ambiguous spoken terms from contextual cues, thereby reducing the dependency on the auditory system. Based on rich linguistic knowledge and powerful reasoning ability of LLMs, this chapter presents recent studies in using LLMs for generative error correction (GER) in ASR to improve recognition results.

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

Large Language Models Meet Speech Recognition

  • Pin-Yu Chen,
  • Sijia Liu

摘要

Advancements in deep neural networks have allowed automatic speech recognition (ASR) systems to attain human parity on several publicly available clean speech datasets. However, even state-of-the-art ASR systems experience performance degradation when confronted with adverse conditions, as a well-trained acoustic model is sensitive to variations in the speech domain, e.g., background noise. On the other hand, linguistic knowledge in large language models (LLMs) can be used to infer the meaning of ambiguous spoken terms from contextual cues, thereby reducing the dependency on the auditory system. Based on rich linguistic knowledge and powerful reasoning ability of LLMs, this chapter presents recent studies in using LLMs for generative error correction (GER) in ASR to improve recognition results.