<p>Unsupervised Domain Adaptation (UDA) aims to transfer the knowledge from source domain to target domain, which always struggles with severe domain shift between the data. The recent progress on visual-language models (VLMs) have provided a promising way to address UDA, leveraging the knowledge from text for more guided adaptation. However, directly deploying such models on downstream UDA tasks with conventional prompt learning can be challenging, which neglects the diversity of visual samples and can cause mis-alignment between modalities, thus lacking flexibility to adapt both modalities dynamically and limiting the cross-domain knowledge transfer. In this paper, we propose an innovative domain-invariant prompt learning method to align the prompts from different modalities. Specifically, we first introduce a hybrid-modality guided prompting module that leverages the prompted multi-modal representation to synergistically help uni-modal learning, thus mutually aligning visual and textual embeddings. We also take advantage of the the category-wise attributes derived from Large Language Model (LLM) to incorporate fine-grained semantic knowledge into prompt learning, ensuring better discrimination among different classes. Besides, to further minimize domain discrepancy, we propose to fuse the textual prototypes with the visual prototypes from each domain, thus to make the input attend to overall domain distribution, which effectively integrates self-enhanced and cross-domain features into the model prediction. With our framework, the two modalities can be mutually promoted to better enhance the adaptation of VLMs for UDA. Experiments on several different benchmarks demonstrate the superiority of our method over previous approaches.</p>

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

Multi-modal Prompt Alignment with Fine-grained LLM Knowledge for Unsupervised Domain Adaptation

  • Bowei Xing,
  • Xianghua Ying,
  • Ruibin Wang,
  • Ruohao Guo

摘要

Unsupervised Domain Adaptation (UDA) aims to transfer the knowledge from source domain to target domain, which always struggles with severe domain shift between the data. The recent progress on visual-language models (VLMs) have provided a promising way to address UDA, leveraging the knowledge from text for more guided adaptation. However, directly deploying such models on downstream UDA tasks with conventional prompt learning can be challenging, which neglects the diversity of visual samples and can cause mis-alignment between modalities, thus lacking flexibility to adapt both modalities dynamically and limiting the cross-domain knowledge transfer. In this paper, we propose an innovative domain-invariant prompt learning method to align the prompts from different modalities. Specifically, we first introduce a hybrid-modality guided prompting module that leverages the prompted multi-modal representation to synergistically help uni-modal learning, thus mutually aligning visual and textual embeddings. We also take advantage of the the category-wise attributes derived from Large Language Model (LLM) to incorporate fine-grained semantic knowledge into prompt learning, ensuring better discrimination among different classes. Besides, to further minimize domain discrepancy, we propose to fuse the textual prototypes with the visual prototypes from each domain, thus to make the input attend to overall domain distribution, which effectively integrates self-enhanced and cross-domain features into the model prediction. With our framework, the two modalities can be mutually promoted to better enhance the adaptation of VLMs for UDA. Experiments on several different benchmarks demonstrate the superiority of our method over previous approaches.