This paper introduces BiGuidedPrompt, a novel prompt generation method that employs a dynamic approach incorporating bidirectional guidance between visual and textual modalities. Unlike traditional methods, such as MaPLe, which rely on static prompts, BiGuidedPrompt utilizes a Prompt Generator that dynamically generates tailored prompts for each image-text pair, enabling more adaptable responses to individual inputs. This tailored prompt generation mechanism dynamically bridges the two modalities, fostering a synergistic encoding process that enhances model generalization across diverse scenarios. Experimental results show that BiGuidedPrompt surpasses existing static prompt models, with mean accuracy improvements of 0.38%, 0.12%, and 0.24% on base, novel, and harmonic-mean metrics, respectively, across 11 datasets. Our contributions include the introduction of a novel Prompt Generator that leverages cross-modal features to create dynamic prompts, thereby improving the alignment of visual and textual branches. Additionally, we show that the logits generated by the Prompt Generator during training yield superior performance compared to those produced by static prompt approaches. These innovations pave the way for future research on more advanced dynamic prompt mechanisms and their applications in multimodal learning tasks.

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

BiGuidedPrompt: Dynamic Bidirectional Guided Multimodal Prompt Learning

  • Jiacheng Zhong,
  • Xinguo Zhang,
  • Bingxue Zhang,
  • Jiasong Wu

摘要

This paper introduces BiGuidedPrompt, a novel prompt generation method that employs a dynamic approach incorporating bidirectional guidance between visual and textual modalities. Unlike traditional methods, such as MaPLe, which rely on static prompts, BiGuidedPrompt utilizes a Prompt Generator that dynamically generates tailored prompts for each image-text pair, enabling more adaptable responses to individual inputs. This tailored prompt generation mechanism dynamically bridges the two modalities, fostering a synergistic encoding process that enhances model generalization across diverse scenarios. Experimental results show that BiGuidedPrompt surpasses existing static prompt models, with mean accuracy improvements of 0.38%, 0.12%, and 0.24% on base, novel, and harmonic-mean metrics, respectively, across 11 datasets. Our contributions include the introduction of a novel Prompt Generator that leverages cross-modal features to create dynamic prompts, thereby improving the alignment of visual and textual branches. Additionally, we show that the logits generated by the Prompt Generator during training yield superior performance compared to those produced by static prompt approaches. These innovations pave the way for future research on more advanced dynamic prompt mechanisms and their applications in multimodal learning tasks.