Recent advancements in Multimodal Large Language Models (MLLMs) have enabled to process diverse input modalities, leading to significantly better understanding of multimedia contents. However, understanding videos is still difficult and even the latest models often create hallucinations. This study introduces a novel method to address event-level hallucinations in MLLMs with a special focus on inferring temporal information of events occurring in an input video. It targets event-related information from both the text query and the video content to enhance MLLMs’ response. Specifically, our method first decomposes these event queries into iconic actions, and then identifies the timestamps of these actions by utilizing external multi-modal models such as CLIP and BLIP2. Experiments using the Charades-STA dataset show that the method decreases the number of hallucinations and improves the MLLM’s responses. We also introduce a quantifiable approach to access these models’ performance in understanding and responding to time-related queries. We designed two question-and-answer tasks to measure response hallucinations in terms of detailed timestamps and the order of time events, respectively. After using our method, the error rates of MLLM’s responses in these two tasks decreased by 39.7% and 36.1%, respectively.

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

Temporal Insight Enhancement: Mitigating Temporal Hallucination in Video Understanding by Multimodal Large Language Models

  • Li Sun,
  • Liuan Wang,
  • Jun Sun,
  • Takayuki Okatani

摘要

Recent advancements in Multimodal Large Language Models (MLLMs) have enabled to process diverse input modalities, leading to significantly better understanding of multimedia contents. However, understanding videos is still difficult and even the latest models often create hallucinations. This study introduces a novel method to address event-level hallucinations in MLLMs with a special focus on inferring temporal information of events occurring in an input video. It targets event-related information from both the text query and the video content to enhance MLLMs’ response. Specifically, our method first decomposes these event queries into iconic actions, and then identifies the timestamps of these actions by utilizing external multi-modal models such as CLIP and BLIP2. Experiments using the Charades-STA dataset show that the method decreases the number of hallucinations and improves the MLLM’s responses. We also introduce a quantifiable approach to access these models’ performance in understanding and responding to time-related queries. We designed two question-and-answer tasks to measure response hallucinations in terms of detailed timestamps and the order of time events, respectively. After using our method, the error rates of MLLM’s responses in these two tasks decreased by 39.7% and 36.1%, respectively.