<p>Parameter-Efficient Fine-Tuning (PEFT) methods have enabled practical adaptation of Large Language Models (LLMs), yet integrating Mixture-of-Experts (MoE) architectures for multi-task scenarios introduces two critical challenges. First, existing MoE-based PEFT approaches overlook the need to balance commonalities and distinctions among diverse tasks, thereby limiting cross-task transferability. Second, prevailing routing strategies rely on discrete expert selection and gradient estimation techniques, which intermittently undermine the ability to consistently identify optimal routing pathways. To address these issues, we propose MoFL (Mixture of Fused LoRA), a hierarchical MoE architecture that incorporates Convergent Experts and Divergent Experts with fused low-rank parameters to jointly capture shared and task-specific knowledge. Complementing this design, a fully differentiable ensemble routing strategy is introduced, eliminating reliance on gradient estimation while enabling collaborative contributions from all experts during generation. Extensive experiments on the HuggingFace Open LLM benchmark demonstrate that MoFL achieves an average score of 52.45 while updating less than 2% of model parameters and consistently outperforms existing LoRA-based PEFT methods across diverse backbone LLMs. Notably, on the GSM8K mathematical reasoning benchmark, the mean score rises from 33.53 to 38.68, the most substantial improvement among all evaluated tasks.</p>

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

MoFL: fine-tuning LLMs with mixture of fused LoRA experts

  • Shengchuan Lin,
  • Dehong Gao,
  • Yufei Ma,
  • Shuangwei Hu,
  • Dejun Mu

摘要

Parameter-Efficient Fine-Tuning (PEFT) methods have enabled practical adaptation of Large Language Models (LLMs), yet integrating Mixture-of-Experts (MoE) architectures for multi-task scenarios introduces two critical challenges. First, existing MoE-based PEFT approaches overlook the need to balance commonalities and distinctions among diverse tasks, thereby limiting cross-task transferability. Second, prevailing routing strategies rely on discrete expert selection and gradient estimation techniques, which intermittently undermine the ability to consistently identify optimal routing pathways. To address these issues, we propose MoFL (Mixture of Fused LoRA), a hierarchical MoE architecture that incorporates Convergent Experts and Divergent Experts with fused low-rank parameters to jointly capture shared and task-specific knowledge. Complementing this design, a fully differentiable ensemble routing strategy is introduced, eliminating reliance on gradient estimation while enabling collaborative contributions from all experts during generation. Extensive experiments on the HuggingFace Open LLM benchmark demonstrate that MoFL achieves an average score of 52.45 while updating less than 2% of model parameters and consistently outperforms existing LoRA-based PEFT methods across diverse backbone LLMs. Notably, on the GSM8K mathematical reasoning benchmark, the mean score rises from 33.53 to 38.68, the most substantial improvement among all evaluated tasks.