Logical reasoning is a manifestation of advanced human cognitive abilities, which requires in-depth analysis and pattern exploration on the basis of fully understanding the content. Multimodal logical reasoning represents a significant research area within logical reasoning and recent methods focus on exploring the capacity of large language models to tackle this task. However, existing multimodal logic reasoning datasets lack rigorous manual validation on automatic generation and perform a uniform problem structure. To fill this gap, we construct a high-quality Multimodal Logical Reasoning dataset, namely MLRQA, which comprises 4,356 multiple pattern questions written by domain experts. We also employ the multimodal chain of thought (CoT) to encourage the reasoning potential of multimodal large language models (MLLMs). Experimental demonstrate that state-of-the-art GPT-4v performs significantly worse than the human benchmark. Our dataset can also serve as an important benchmark for investigating multimodal logical reasoning in future. The dataset is available at https://github.com/pingzv/MLRQA .

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

MLRQA: A Dataset with Multimodal Logical Reasoning Challenges

  • Jing Xiao,
  • Guijin Lin,
  • Ping Li

摘要

Logical reasoning is a manifestation of advanced human cognitive abilities, which requires in-depth analysis and pattern exploration on the basis of fully understanding the content. Multimodal logical reasoning represents a significant research area within logical reasoning and recent methods focus on exploring the capacity of large language models to tackle this task. However, existing multimodal logic reasoning datasets lack rigorous manual validation on automatic generation and perform a uniform problem structure. To fill this gap, we construct a high-quality Multimodal Logical Reasoning dataset, namely MLRQA, which comprises 4,356 multiple pattern questions written by domain experts. We also employ the multimodal chain of thought (CoT) to encourage the reasoning potential of multimodal large language models (MLLMs). Experimental demonstrate that state-of-the-art GPT-4v performs significantly worse than the human benchmark. Our dataset can also serve as an important benchmark for investigating multimodal logical reasoning in future. The dataset is available at https://github.com/pingzv/MLRQA .