The wide linguistic diversity in Indigenous regions can create critical situations when these people need urgent medical or legal services and no one can identify the language they speak. Even if the institution has a directory of Indigenous interpreters, how do they know which one to contact? As a first step toward addressing this problem, this study explores linguistic identification between spoken Yucatec Maya and Spanish using deep learning models. A dataset of 3,820 audio files was selected, consisting of 2,116 clips in Spanish and 1,704 clips in Yucatec Maya. The audio files were segmented into 3 to 15 s fragments and gender-balanced where possible. RGB spectrograms were extracted from the audio segments for model input. Three deep learning architectures (ResNet-152, ResNet-50, and MobileNetV2) were evaluated to compare their accuracy, precision, recall, and F1 score performance. The results indicate that ResNet-152 achieved the highest accuracy (90.57%) and F1 score (0.9010), demonstrating its effectiveness in language identification tasks. MobileNetV2, although slightly less accurate (87.17%), showed superior efficiency, making it suitable for resource-limited environments. The learning curves indicate the need for more data or to expand the evaluation to other resource-limited languages to improve the generalizability of the results. The study highlights the potential of deep learning for resource-limited language identification and provides insight into the relationship between model performance and computational efficiency. These findings not only contribute to the development of tools for preserving Indigenous languages but can also expedite the search for human interpreters in emergency situations.

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

Bridging the Language Gap: AI-Powered Identification of Yucatec Maya and Spanish for Emergency Response

  • Jean Carlos Buenfil Aguilar,
  • Silvia Fernández-Sabido,
  • Jorge Reyes-Magaña,
  • Luis Basto-Díaz,
  • Luis Fernando Curi-Quintal,
  • Oscar Gerardo Sánchez Siordia

摘要

The wide linguistic diversity in Indigenous regions can create critical situations when these people need urgent medical or legal services and no one can identify the language they speak. Even if the institution has a directory of Indigenous interpreters, how do they know which one to contact? As a first step toward addressing this problem, this study explores linguistic identification between spoken Yucatec Maya and Spanish using deep learning models. A dataset of 3,820 audio files was selected, consisting of 2,116 clips in Spanish and 1,704 clips in Yucatec Maya. The audio files were segmented into 3 to 15 s fragments and gender-balanced where possible. RGB spectrograms were extracted from the audio segments for model input. Three deep learning architectures (ResNet-152, ResNet-50, and MobileNetV2) were evaluated to compare their accuracy, precision, recall, and F1 score performance. The results indicate that ResNet-152 achieved the highest accuracy (90.57%) and F1 score (0.9010), demonstrating its effectiveness in language identification tasks. MobileNetV2, although slightly less accurate (87.17%), showed superior efficiency, making it suitable for resource-limited environments. The learning curves indicate the need for more data or to expand the evaluation to other resource-limited languages to improve the generalizability of the results. The study highlights the potential of deep learning for resource-limited language identification and provides insight into the relationship between model performance and computational efficiency. These findings not only contribute to the development of tools for preserving Indigenous languages but can also expedite the search for human interpreters in emergency situations.