Performant Multilingual Modulated and Multiplexed Memory Distilled Model with Adaptive Activation Ensembles
摘要
Modern Multilingual Large Language Models are built on gigantic networks, and training them in constrained environments is not always feasible. Modern Multilingual Large Language Models are also resource-intensive, and training these networks is time-consuming. PM4DA2E (Performant Multilingual Modulated and Multiplexed Memory Distilled Model with Adaptive Activation Ensembles) outshines by overcoming the challenges of Natural Language Processing Multilingual Large Language Models. PM4DA2E provides an alternate solution that is reliable and better suitable for constraint setups by introducing amendments to the novel transformer-based architecture with distillation, modulation, multiplexing, adapters, memory, and activation ensembles. PM4DA2E, due to its structural transformations, retains 99% of knowledge with 50% reduced network size. PM4DA2E achieves 51% inference speedup and 47% less training time on the XNLI (Cross-lingual Natural Language Inference) dataset / XNLI evaluation metrics (Accuracy and/or F1-Score) compared to competing Multilingual Large Language Network counterparts prevalent in this era.