<p>Large Vision-Language Models (LVLMs) demonstrate strong capabilities but remain susceptible to hallucinations. To address this limitation, we propose reflective instruction tuning, which explicitly trains models to reflect. Rather than producing only a final answer, the model is supervised to generate reflective rationales that justify the correct prediction and explain why plausible alternatives are incorrect. To remain effective for advanced LVLMs deployed in diverse real-world scenarios, reflective supervision must be more fine-grained and broader in domains. We therefore propose REVERIE+ (<b>R</b>efl<b>E</b>cti<b>VE</b> <b>R</b>at<b>I</b>onal<b>E</b>), an extension of REVERIE substantially expanded in domain diversity, task complexity, and annotation richness, tailored to advanced LVLMs. Built on the R1-Onevision data foundation, REVERIE+ broadens domain coverage and increases task difficulty, while improving annotation reliability by (i) leveraging multiple models to mine and label diverse negative answers, and (ii) employing a strong commercial LVLM to generate higher-quality reflective rationales. These changes provide richer negative supervision and more accurate reflection signals, making REVERIE+ better suited for hallucination mitigation in more capable models. Extensive experiments across representative hallucination and general multimodal benchmarks show consistent performance gains, validating that the enriched negative supervision and precise reflection signals of REVERIE+ offer a scalable and effective approach for mitigating hallucinations in advanced LVLMs.</p>

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

REVERIE+: Generalized Reflective Instruction Tuning for Hallucination Mitigation in Advanced VLMs

  • Mingyang Bi,
  • Jinrui Zhang,
  • Xiangchen Wang,
  • Xue Jiang,
  • Yuhang Lu,
  • Peng Wang,
  • Feng Zheng

摘要

Large Vision-Language Models (LVLMs) demonstrate strong capabilities but remain susceptible to hallucinations. To address this limitation, we propose reflective instruction tuning, which explicitly trains models to reflect. Rather than producing only a final answer, the model is supervised to generate reflective rationales that justify the correct prediction and explain why plausible alternatives are incorrect. To remain effective for advanced LVLMs deployed in diverse real-world scenarios, reflective supervision must be more fine-grained and broader in domains. We therefore propose REVERIE+ (ReflEctiVE RatIonalE), an extension of REVERIE substantially expanded in domain diversity, task complexity, and annotation richness, tailored to advanced LVLMs. Built on the R1-Onevision data foundation, REVERIE+ broadens domain coverage and increases task difficulty, while improving annotation reliability by (i) leveraging multiple models to mine and label diverse negative answers, and (ii) employing a strong commercial LVLM to generate higher-quality reflective rationales. These changes provide richer negative supervision and more accurate reflection signals, making REVERIE+ better suited for hallucination mitigation in more capable models. Extensive experiments across representative hallucination and general multimodal benchmarks show consistent performance gains, validating that the enriched negative supervision and precise reflection signals of REVERIE+ offer a scalable and effective approach for mitigating hallucinations in advanced LVLMs.