<p>Recently, fine-tuning large pre-trained models has yielded promising results in few-shot learning regime. However, their ability to generalize on two-dimensional Out-of-Distribution (OoD) data, characterized by correlation shift and diversity shift, remains largely unexplored. Recent studies demonstrate that even with a vast amount of data, few methods can beat the empirical risk minimization method (ERM) simultaneously on the two-dimensional OoD generalization. Consequently, OoD generalization in the few-shot regime emerges as a challenging area in model generalization research, where the OoD test performance suffers from the synergies of both data distribution shifts and overfitting few-shot samples. In this paper, utilizing informative natural language supervision, we investigate a novel Bayesian cross-modal image-text alignment learning method (Bayes-CAL) to address this issue, where only the texts representations are fine-tuned via image-text alignment in a domain-invariant manner under the proposed regularization. The Bayesian approach is essentially introduced to avoid overfitting the base classes observed in the training process and improve generalization to unseen classes. Under mild assumptions, we have theoretically demonstrated our Bayes-CAL method can achieve lower generalization errors under distribution shifts, compared to previous methods. To validate the effectiveness of the proposed Bayes-CAL, in addition to the extensive experiments on the image classification task, we also verify the OoD generalization performance on the object detection and instance segmentation tasks, addressing the gap in previous works on the challenging few-shot OoD generalization. Compared with recent CLIP-based methods, Bayes-CAL achieved state-of-the-art OoD generalization performance on these tasks with a large margin. More stable generalization performances on unseen classes of the Bayes-CAL are further demonstrated in the experiments.</p>

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

Bayes-CAL: Robust Cross-Modal Alignment by Bayesian Approach for Few-Shot OoD Generalization

  • Lin Zhu,
  • Weihan Yin,
  • Fan Wu,
  • Qinying Gu,
  • Xinbing Wang,
  • Chenghu Zhou,
  • Nanyang Ye

摘要

Recently, fine-tuning large pre-trained models has yielded promising results in few-shot learning regime. However, their ability to generalize on two-dimensional Out-of-Distribution (OoD) data, characterized by correlation shift and diversity shift, remains largely unexplored. Recent studies demonstrate that even with a vast amount of data, few methods can beat the empirical risk minimization method (ERM) simultaneously on the two-dimensional OoD generalization. Consequently, OoD generalization in the few-shot regime emerges as a challenging area in model generalization research, where the OoD test performance suffers from the synergies of both data distribution shifts and overfitting few-shot samples. In this paper, utilizing informative natural language supervision, we investigate a novel Bayesian cross-modal image-text alignment learning method (Bayes-CAL) to address this issue, where only the texts representations are fine-tuned via image-text alignment in a domain-invariant manner under the proposed regularization. The Bayesian approach is essentially introduced to avoid overfitting the base classes observed in the training process and improve generalization to unseen classes. Under mild assumptions, we have theoretically demonstrated our Bayes-CAL method can achieve lower generalization errors under distribution shifts, compared to previous methods. To validate the effectiveness of the proposed Bayes-CAL, in addition to the extensive experiments on the image classification task, we also verify the OoD generalization performance on the object detection and instance segmentation tasks, addressing the gap in previous works on the challenging few-shot OoD generalization. Compared with recent CLIP-based methods, Bayes-CAL achieved state-of-the-art OoD generalization performance on these tasks with a large margin. More stable generalization performances on unseen classes of the Bayes-CAL are further demonstrated in the experiments.