Long-tailed image classification poses significant challenges due to severe data imbalance. The scarcity of samples in tail classes results in limited visual information, making it difficult for models to extract meaningful features from those underrepresented categories. To address this issue, we propose a novel approach called Multi-Modal Information Enhancement, which leverages both textual and visual modalities to enrich feature representations and enhance classification performance under long-tailed distributions. Specifically, we introduce Text Enhancement with Semantic Expansion to generate rich and diverse textual descriptions from both feature and scenario dimensions, offering valuable semantic cues for tail classes. Meanwhile, we develop a visual augmentation strategy, named Minority-category Compensation Mixup, to promote the participation of tail classes during training via dual sampling and dynamic label adjustment. Experiments on CIFAR100-LT, CIFAR10-LT, and ImageNet-LT demonstrate that our proposed method effectively alleviates data scarcity and significantly improves recognition performance for tail classes, providing a robust and generalizable solution to long-tailed image classification.

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

Multi-modal Information Enhancement for Long-Tailed Recognition

  • Shengnan Fan,
  • Zhilei Chai,
  • Qin Wu,
  • Xiangyu Cheng,
  • Yuying Pan

摘要

Long-tailed image classification poses significant challenges due to severe data imbalance. The scarcity of samples in tail classes results in limited visual information, making it difficult for models to extract meaningful features from those underrepresented categories. To address this issue, we propose a novel approach called Multi-Modal Information Enhancement, which leverages both textual and visual modalities to enrich feature representations and enhance classification performance under long-tailed distributions. Specifically, we introduce Text Enhancement with Semantic Expansion to generate rich and diverse textual descriptions from both feature and scenario dimensions, offering valuable semantic cues for tail classes. Meanwhile, we develop a visual augmentation strategy, named Minority-category Compensation Mixup, to promote the participation of tail classes during training via dual sampling and dynamic label adjustment. Experiments on CIFAR100-LT, CIFAR10-LT, and ImageNet-LT demonstrate that our proposed method effectively alleviates data scarcity and significantly improves recognition performance for tail classes, providing a robust and generalizable solution to long-tailed image classification.