Multi-modal Information Enhancement for Long-Tailed Recognition
摘要
Long-tailed image classification poses significant challenges due to severe data imbalance. The scarcity of samples in tail classes results in limited visual information, making it difficult for models to extract meaningful features from those underrepresented categories. To address this issue, we propose a novel approach called Multi-Modal Information Enhancement, which leverages both textual and visual modalities to enrich feature representations and enhance classification performance under long-tailed distributions. Specifically, we introduce Text Enhancement with Semantic Expansion to generate rich and diverse textual descriptions from both feature and scenario dimensions, offering valuable semantic cues for tail classes. Meanwhile, we develop a visual augmentation strategy, named Minority-category Compensation Mixup, to promote the participation of tail classes during training via dual sampling and dynamic label adjustment. Experiments on CIFAR100-LT, CIFAR10-LT, and ImageNet-LT demonstrate that our proposed method effectively alleviates data scarcity and significantly improves recognition performance for tail classes, providing a robust and generalizable solution to long-tailed image classification.