The new focus of the international community to return to Moon (Smith et al. in 2020 IEEE aerospace conference, pp. 1–10, 2020) and Mars has sparked the need for new ideas of future manned missions. Unmanned vehicles are being considered to aid astronaut activities on extraterrestrial surfaces (Allak et al. in Astrobiology 20:1321–1337, 2020). This paper is exploring the concept of controlling unmanned ground vehicles (UGV) with voice commands by evaluating existing technologies; specifically, DeepSpeech (DS) (Hannun et al. in Deep speech: Scaling up end-to-end speech recognition, 2014) and its pre-trained v0.6.0 model. The evaluation entails the capability to transcribe spoken English and distinguish edge cases of discrete commands. It will be discussed in depth how the model behaves for various categories of edge cases, including anagrams, rhymes, default commands, and coincidental inclusions. Moreover, technologies like natural language processing and machine learning will be discussed regarding functionality and benefits for command recognition in post-processing as enhancements to plain speech to text conversion for building a voice command interface for an unmanned ground vehicle.

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

Benchmarking Speech-to-Text Systems for Space-Related Applications

  • Tobias Kolb,
  • Falk Schiffner

摘要

The new focus of the international community to return to Moon (Smith et al. in 2020 IEEE aerospace conference, pp. 1–10, 2020) and Mars has sparked the need for new ideas of future manned missions. Unmanned vehicles are being considered to aid astronaut activities on extraterrestrial surfaces (Allak et al. in Astrobiology 20:1321–1337, 2020). This paper is exploring the concept of controlling unmanned ground vehicles (UGV) with voice commands by evaluating existing technologies; specifically, DeepSpeech (DS) (Hannun et al. in Deep speech: Scaling up end-to-end speech recognition, 2014) and its pre-trained v0.6.0 model. The evaluation entails the capability to transcribe spoken English and distinguish edge cases of discrete commands. It will be discussed in depth how the model behaves for various categories of edge cases, including anagrams, rhymes, default commands, and coincidental inclusions. Moreover, technologies like natural language processing and machine learning will be discussed regarding functionality and benefits for command recognition in post-processing as enhancements to plain speech to text conversion for building a voice command interface for an unmanned ground vehicle.