A Meta-thinking Approach to Mitigating Linguistic Sycophancy in Vision-Language Models
摘要
In this paper, we address the critical but underexplored issue of linguistic sycophancy in large vision-language models (LVLMs), where models may alter accurate answers based on incorrect user feedback. Despite extensive studies on sycophantic behavior in large language models (LLMs), its effects on LVLMs remain largely unexamined. We present the first large-scale study on this issue, evaluating 11 open-source and 3 closed-source state-of-the-art models. Our analysis uncovers significant performance declines in models like GPT-4o on complex visual tasks, while LLaVA-RLHF, an open-source model trained with reinforcement learning from human feedback, shows resilience comparable to top closed-source models. We propose a training-free mitigation approach based on meta-thinking, utilizing contextual prompts to encourage models to maintain accuracy and stability despite user feedback. This method significantly improves model reliability and accuracy across various LVLMs, with notable gains in high-performance models, and offers insights into the biases driving sycophantic behavior.