Bilingual Speech Translation Between Tamil and Telugu with Elimination of Dependency on Any Intermediate Languages
摘要
Usage of English as an intermediate language in translating one language to another is prominent, in most of the existing works. Many languages share a common parent language, and when such languages use English as an intermediate language for translation, there can be contextual differences, which could be avoided if there existed a direct source language to target language translation, instead of an external language dependency. Translation of Tamil speech to Telugu speech is discussed here. There are two main approaches: direct translation approach, and cascaded translation approach. A direct speech translation approach using discrete units is compared with a cascaded speech translation system approach that integrates Automatic Speech Recognition, Text-to-Text translation, and Text-to-Speech modules. Several combinations of the cascaded speech translation approach are compared with the direct speech translation approach by training and testing with Few-shot Learning Evaluation of Universal Representations of Speech (FLEURS) dataset. Bilingual Evaluation Understudy (BLEU) and Perceptual Evaluation of Speech Quality (PESQ) evaluatory metrics show that direct Speech-to-Speech Translation is better than some of the cascaded models in the quality of speech translated, but all cascaded Speech-to-Speech Translation approaches are better than the direct Speech-to-Speech Translation approach in preserving the context. It can be concluded that when the right mix of models is cascaded, the translation of Tamil to Telugu becomes better than using a direct Speech-to-Speech Translation.