Usage of English as an intermediate language in translating one language to another is prominent, in most of the existing works. Many languages share a common parent language, and when such languages use English as an intermediate language for translation, there can be contextual differences, which could be avoided if there existed a direct source language to target language translation, instead of an external language dependency. Translation of Tamil speech to Telugu speech is discussed here. There are two main approaches: direct translation approach, and cascaded translation approach. A direct speech translation approach using discrete units is compared with a cascaded speech translation system approach that integrates Automatic Speech Recognition, Text-to-Text translation, and Text-to-Speech modules. Several combinations of the cascaded speech translation approach are compared with the direct speech translation approach by training and testing with Few-shot Learning Evaluation of Universal Representations of Speech (FLEURS) dataset. Bilingual Evaluation Understudy (BLEU) and Perceptual Evaluation of Speech Quality (PESQ) evaluatory metrics show that direct Speech-to-Speech Translation is better than some of the cascaded models in the quality of speech translated, but all cascaded Speech-to-Speech Translation approaches are better than the direct Speech-to-Speech Translation approach in preserving the context. It can be concluded that when the right mix of models is cascaded, the translation of Tamil to Telugu becomes better than using a direct Speech-to-Speech Translation.

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

Bilingual Speech Translation Between Tamil and Telugu with Elimination of Dependency on Any Intermediate Languages

  • R. Kaarthik,
  • S. Rithuh Subhakkrith,
  • Yasaswini Dharmavarapu,
  • A. Vignesh Kumar,
  • Dhanya M. Dhanalakshmy,
  • N. G. Karthikeyan

摘要

Usage of English as an intermediate language in translating one language to another is prominent, in most of the existing works. Many languages share a common parent language, and when such languages use English as an intermediate language for translation, there can be contextual differences, which could be avoided if there existed a direct source language to target language translation, instead of an external language dependency. Translation of Tamil speech to Telugu speech is discussed here. There are two main approaches: direct translation approach, and cascaded translation approach. A direct speech translation approach using discrete units is compared with a cascaded speech translation system approach that integrates Automatic Speech Recognition, Text-to-Text translation, and Text-to-Speech modules. Several combinations of the cascaded speech translation approach are compared with the direct speech translation approach by training and testing with Few-shot Learning Evaluation of Universal Representations of Speech (FLEURS) dataset. Bilingual Evaluation Understudy (BLEU) and Perceptual Evaluation of Speech Quality (PESQ) evaluatory metrics show that direct Speech-to-Speech Translation is better than some of the cascaded models in the quality of speech translated, but all cascaded Speech-to-Speech Translation approaches are better than the direct Speech-to-Speech Translation approach in preserving the context. It can be concluded that when the right mix of models is cascaded, the translation of Tamil to Telugu becomes better than using a direct Speech-to-Speech Translation.