Large Language Models Meet Speech Recognition
摘要
Advancements in deep neural networks have allowed automatic speech recognition (ASR) systems to attain human parity on several publicly available clean speech datasets. However, even state-of-the-art ASR systems experience performance degradation when confronted with adverse conditions, as a well-trained acoustic model is sensitive to variations in the speech domain, e.g., background noise. On the other hand, linguistic knowledge in large language models (LLMs) can be used to infer the meaning of ambiguous spoken terms from contextual cues, thereby reducing the dependency on the auditory system. Based on rich linguistic knowledge and powerful reasoning ability of LLMs, this chapter presents recent studies in using LLMs for generative error correction (GER) in ASR to improve recognition results.