An Overview of Automatic Speech Recognition Based on Deep Learning and Bio–Signal Sensors
摘要
The process of producing speech is complex and includes a number of biosignals in addition to acoustics. In order to overcome the limitations of conventional speech processing in particular and to gain a better understanding of the process of speech creation in general, these biosignals can be used. They originate from the articulators, the movement of the articulator muscles, the connections within the brain, and the brain itself. With an emphasis on speech production, recognition, and volitional control, we discuss artificial mouth techniques in this review that make use of a variety of sensors, including gyros, images, 3-axial magnetic sensors, electromyography as EMG, electroencephalography as EEG, electropalatography as EPG, electromagnetic articulography as EMA, permanent magnet articulography as PMA, and articulator electromyography. Before classifying them into taxonomy, we evaluate the flow of several voice recognition-related deep learning technologies, including visual speech recognition and silent speech interface. We conclude by talking about ways to address the communication issues that persons with speech impairments face as well as upcoming deep learning research.