\(E^3\) : Optimizing Language Model Training for Translation via Enhancing Efficiency and Effectiveness
摘要
In the field of Natural Language Processing (NLP), Large-scale Language Models (LLMs) have demonstrated exceptional capabilities across a variety of tasks, including question answering, classification, and particularly natural language understanding. The integration of neural machine translation with LLMs presents significant potential, transforming the paradigms of cross-lingual communication and information exchange. This study investigates the foundational aspects of LLMs’ translation abilities and identifies effective training methodologies to equip them with multilingual capacities. We specifically explore the optimal timing for introducing translation capabilities to LLMs via supervised tasks, considering the inherent bilingual nature of machine translation. Key questions explored include whether it is more beneficial to integrate multiple languages during the pre-training or supervised fine-tuning (SFT) stages, how variations in language ratios influence LLMs’ translation abilities, and whether longer or shorter texts are more effective for training these models. This research conducts a thorough investigation by training multiple LLMs from scratch with parameter scales in the billions and enhances the robustness of our findings by upgrading the language capabilities of pre-trained open-source models with parameter scales reaching tens of billions. The aim is to provide a detailed analysis that elucidates the complexities of augmenting machine translation capabilities within LLMs.