Long-tailed recognition (LTR) remains challenging due to data imbalance and distribution bias. Existing methods may overfit tail classes or degrade head class performance, and pre-trained models such as Contrastive Language-Image Pre-Training (CLIP) exhibit limited adaptability. To address these issues, we propose MSCF: a framework integrating Multi-Sampling Classifiers with a Frozen CLIP backbone. Specifically, MSCF initially employs a frozen CLIP visual backbone to extract image features, thereby avoiding the computational costs and overfitting risks associated with fine-tuning. To address the characteristics of LTR, this study constructs multi-sampling classifiers designed to learn and adapt to the feature representations of different data distributions. For handling LTR, MSCF incorporates three distinct sampling strategies: original sampling, balanced sampling, and inverse sampling. The final classification result is obtained by fusing the outputs of multi-sampling classifiers. Experiments on three LTR benchmarks confirm that MSCF captures data characteristics across multiple sampling strategies without backbone fine-tuning. Moreover, it outperforms conventional approaches. Visualization and ablation studies reveal that MSCF effectively disentangles inherent semantic features from data bias, providing new insights into adapting pre-trained models for long-tailed tasks.

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

Long-Tailed Recognition via Multi-sampling Classifier Fusion with Frozen CLIP

  • Yilou Zhang,
  • Yunjie Liu,
  • Yuan Zhao,
  • Chengkun Wu

摘要

Long-tailed recognition (LTR) remains challenging due to data imbalance and distribution bias. Existing methods may overfit tail classes or degrade head class performance, and pre-trained models such as Contrastive Language-Image Pre-Training (CLIP) exhibit limited adaptability. To address these issues, we propose MSCF: a framework integrating Multi-Sampling Classifiers with a Frozen CLIP backbone. Specifically, MSCF initially employs a frozen CLIP visual backbone to extract image features, thereby avoiding the computational costs and overfitting risks associated with fine-tuning. To address the characteristics of LTR, this study constructs multi-sampling classifiers designed to learn and adapt to the feature representations of different data distributions. For handling LTR, MSCF incorporates three distinct sampling strategies: original sampling, balanced sampling, and inverse sampling. The final classification result is obtained by fusing the outputs of multi-sampling classifiers. Experiments on three LTR benchmarks confirm that MSCF captures data characteristics across multiple sampling strategies without backbone fine-tuning. Moreover, it outperforms conventional approaches. Visualization and ablation studies reveal that MSCF effectively disentangles inherent semantic features from data bias, providing new insights into adapting pre-trained models for long-tailed tasks.