Estimating Tongue Movements from Speech Using Deep Learning and Machine Learning Techniques
摘要
This study presents an innovative method to determine tongue movement during speech. This is accomplished by using ultrasound tongue imaging (UTI) from the TaL80 dataset. Our methodology uses advanced deep learning and machine learning techniques. First, a feature extractor is built by using transfer learning using ResNet50 to extract features. Moreover, we use the Extra Tree Classifier to select features effectively. Finally, we use Support Vector Machines (SVMs) to perform classification tasks. Our experiments indicate the effectiveness of the recommended technique in estimating tongue movements with a high accuracy of 100%, and also the validity of our results is verified by employing evaluation measures such as F1-score, recall, and precision. These measures achieved a 100% success rate by examining multiple cases, including one without feature selection and SMOTE technique, and another with feature selection and SMOTE technique. This allowed us to select the most optimal solution. The proposed technique shows great potential for use in speech therapy, language learning, assistive communication devices, and other related domains, opening up new possibilities for future research.