In this paper, we address the critical but underexplored issue of linguistic sycophancy in large vision-language models (LVLMs), where models may alter accurate answers based on incorrect user feedback. Despite extensive studies on sycophantic behavior in large language models (LLMs), its effects on LVLMs remain largely unexamined. We present the first large-scale study on this issue, evaluating 11 open-source and 3 closed-source state-of-the-art models. Our analysis uncovers significant performance declines in models like GPT-4o on complex visual tasks, while LLaVA-RLHF, an open-source model trained with reinforcement learning from human feedback, shows resilience comparable to top closed-source models. We propose a training-free mitigation approach based on meta-thinking, utilizing contextual prompts to encourage models to maintain accuracy and stability despite user feedback. This method significantly improves model reliability and accuracy across various LVLMs, with notable gains in high-performance models, and offers insights into the biases driving sycophantic behavior.

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

A Meta-thinking Approach to Mitigating Linguistic Sycophancy in Vision-Language Models

  • Chinh Hoang,
  • Nathan Roberts,
  • Mohammad Rashedul Hasan

摘要

In this paper, we address the critical but underexplored issue of linguistic sycophancy in large vision-language models (LVLMs), where models may alter accurate answers based on incorrect user feedback. Despite extensive studies on sycophantic behavior in large language models (LLMs), its effects on LVLMs remain largely unexamined. We present the first large-scale study on this issue, evaluating 11 open-source and 3 closed-source state-of-the-art models. Our analysis uncovers significant performance declines in models like GPT-4o on complex visual tasks, while LLaVA-RLHF, an open-source model trained with reinforcement learning from human feedback, shows resilience comparable to top closed-source models. We propose a training-free mitigation approach based on meta-thinking, utilizing contextual prompts to encourage models to maintain accuracy and stability despite user feedback. This method significantly improves model reliability and accuracy across various LVLMs, with notable gains in high-performance models, and offers insights into the biases driving sycophantic behavior.