<p>Automatic Speech Recognition (ASR) plays a crucial role in human-machine interaction, particularly in industrial environments. This work presents an Arabic ASR system designed for industrial command recognition, utilizing a stochastic Hidden Markov Model (HMM) approach. While deep learning-based ASR models dominate the field, their dependence on extensive labeled datasets limits their applicability to under-resourced dialects. To address this issue, an efficient and adaptable alternative is proposed, enabling real-time deployment on embedded platforms. The main contributions include the development of an optimized ASR system suitable for low-cost embedded hardware, an in-depth phonetic and syllabic analysis to enhance recognition accuracy, and a comprehensive evaluation of computational efficiency on a Raspberry Pi 4. The system achieves a recognition accuracy of 92.05% using a 3-HMM architecture with 16-Gaussian Mixture Models, demonstrating its effectiveness for industrial command applications. Despite hardware constraints, real-time performance remains viable, with execution times ranging from 88.39&#xa0;ms on a Raspberry Pi to 32.63&#xa0;ms on a laptop for complex commands. Resource utilization is also well-managed, with CPU usage increasing by a factor of 2.5–3, execution time extending by 2–3 times, and memory consumption remaining below 4.8 MB on the Raspberry Pi. These results highlight the feasibility of deploying Arabic ASR in industrial settings, advancing speech recognition in resource-constrained environments.</p>

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

Embedded speech recognition for Arabic language in industrial command

  • Naouar Laaidi,
  • Abderrahim Ezzine,
  • Hassan Satori

摘要

Automatic Speech Recognition (ASR) plays a crucial role in human-machine interaction, particularly in industrial environments. This work presents an Arabic ASR system designed for industrial command recognition, utilizing a stochastic Hidden Markov Model (HMM) approach. While deep learning-based ASR models dominate the field, their dependence on extensive labeled datasets limits their applicability to under-resourced dialects. To address this issue, an efficient and adaptable alternative is proposed, enabling real-time deployment on embedded platforms. The main contributions include the development of an optimized ASR system suitable for low-cost embedded hardware, an in-depth phonetic and syllabic analysis to enhance recognition accuracy, and a comprehensive evaluation of computational efficiency on a Raspberry Pi 4. The system achieves a recognition accuracy of 92.05% using a 3-HMM architecture with 16-Gaussian Mixture Models, demonstrating its effectiveness for industrial command applications. Despite hardware constraints, real-time performance remains viable, with execution times ranging from 88.39 ms on a Raspberry Pi to 32.63 ms on a laptop for complex commands. Resource utilization is also well-managed, with CPU usage increasing by a factor of 2.5–3, execution time extending by 2–3 times, and memory consumption remaining below 4.8 MB on the Raspberry Pi. These results highlight the feasibility of deploying Arabic ASR in industrial settings, advancing speech recognition in resource-constrained environments.