Continual learning systems must handle sequential tasks while preserving previously acquired knowledge-a challenge often hindered by catastrophic forgetting. We propose TMMP (Top Maximum Magnitude Parameters), which leverages task-specific low-rank adapters (LoRA) within a Mixture of Experts (MoE) framework alongside a strategic model merging strategy. TMMP creates dedicated expert modules for each sequential task and employs a parameter fusion technique that selectively merges weights based on magnitude significance. This approach minimizes interference between tasks while maximizing knowledge retention. We evaluate TMMP on class-incremental learning scenarios using CIFAR100 and TinyImageNet benchmark datasets. Our experimental results demonstrate that TMMP outperforms state-of-the-art methods in accuracy while significantly reducing forgetting, particularly when handling long task sequences. The parameter-efficient MoE architecture combined with strategic model merging mechanism provides an effective solution for continual learning in vision-language models.

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

TMMP: Efficient Continual Learning via Mixture of LoRA Experts and Top Maximum Magnitude Parameter Fusion

  • Quynh-Trang Pham Thi,
  • Le-Huy Pham,
  • Dinh-Dat Nguyen,
  • Tri-Thanh Nguyen,
  • Thanh Hai Dang

摘要

Continual learning systems must handle sequential tasks while preserving previously acquired knowledge-a challenge often hindered by catastrophic forgetting. We propose TMMP (Top Maximum Magnitude Parameters), which leverages task-specific low-rank adapters (LoRA) within a Mixture of Experts (MoE) framework alongside a strategic model merging strategy. TMMP creates dedicated expert modules for each sequential task and employs a parameter fusion technique that selectively merges weights based on magnitude significance. This approach minimizes interference between tasks while maximizing knowledge retention. We evaluate TMMP on class-incremental learning scenarios using CIFAR100 and TinyImageNet benchmark datasets. Our experimental results demonstrate that TMMP outperforms state-of-the-art methods in accuracy while significantly reducing forgetting, particularly when handling long task sequences. The parameter-efficient MoE architecture combined with strategic model merging mechanism provides an effective solution for continual learning in vision-language models.