<p>Multi-modal multi-label emotion recognition (MMER) aims to identify multiple co-occurring emotions from heterogeneous modalities. Existing methods tend to emphasize either modality complementarity or modality-specific information, while rarely achieving a balanced integration of both. Moreover, dependencies between modalities and emotion labels remain insufficiently explored, and the challenge of label imbalance in MMER has received little attention. To address these challenges, we propose COLIN (COmplementary and competitive baLanced learnIng Network), a unified framework designed to simultaneously balance modality complementarity and specificity while alleviating learning imbalance. Specifically, COLIN incorporates a Mixture-of-Experts (MoE) mechanism with attention to capture label dependencies and optimize multi-modal feature representations. It further adopts a dual-branch architecture that integrates modality-specific feature reconstruction with multi-level fusion, ensuring balanced utilization of information across different modalities. In addition, an improved loss function is designed to mitigate label imbalance and enhance overall model robustness. Extensive experiments on two benchmark datasets, CMU-MOSEI and <InlineEquation ID="IEq1"> <EquationSource Format="TEX">\(\text {M}^3\text {ED}\)</EquationSource> <EquationSource Format="MATHML"><math> <mrow> <msup> <mtext>M</mtext> <mn>3</mn> </msup> <mtext>ED</mtext> </mrow> </math></EquationSource> </InlineEquation>, demonstrate that COLIN consistently outperforms state-of-the-art methods, validating its effectiveness in both modality balancing and robust multi-label learning.</p>

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

COLIN: complementary and competitive balanced learning network for multi-modal multi-label emotion recognition

  • Xiaoyu Liu,
  • Ting Wang,
  • Aixiang Cui,
  • Xiaowen Zhang

摘要

Multi-modal multi-label emotion recognition (MMER) aims to identify multiple co-occurring emotions from heterogeneous modalities. Existing methods tend to emphasize either modality complementarity or modality-specific information, while rarely achieving a balanced integration of both. Moreover, dependencies between modalities and emotion labels remain insufficiently explored, and the challenge of label imbalance in MMER has received little attention. To address these challenges, we propose COLIN (COmplementary and competitive baLanced learnIng Network), a unified framework designed to simultaneously balance modality complementarity and specificity while alleviating learning imbalance. Specifically, COLIN incorporates a Mixture-of-Experts (MoE) mechanism with attention to capture label dependencies and optimize multi-modal feature representations. It further adopts a dual-branch architecture that integrates modality-specific feature reconstruction with multi-level fusion, ensuring balanced utilization of information across different modalities. In addition, an improved loss function is designed to mitigate label imbalance and enhance overall model robustness. Extensive experiments on two benchmark datasets, CMU-MOSEI and \(\text {M}^3\text {ED}\) M 3 ED , demonstrate that COLIN consistently outperforms state-of-the-art methods, validating its effectiveness in both modality balancing and robust multi-label learning.