Can Large Vision-Language Models Conduct Advanced Semantic Understanding for Counter-Commonsense Scenarios?
摘要
Existing research on counter-commonsense (CC) reasoning in large vision-language models primarily focuses on simple attribute-based reasoning, with limited evaluation of advanced semantic understanding. To address this gap, we introduce the Counter-Commonsense Benchmark (CC-Bench), a novel framework assessing advanced semantic understanding and logical consistency in multi-modal models. CC-Bench comprises 1,051 CC examples and includes three tasks: CC Judgment (CC-J), CC Object Judgment (CC-OJ), and CC Scene Classification (CC-SC). It also introduces two logical metrics to evaluate models’ logical consistency during reasoning. Experiments on nine state-of-the-art multi-modal large language models (MLLMs) revealed that top accuracy was below 65% on both CC-OJ and CC-SC tasks, indicating insufficient advanced semantic understanding. Logical consistency metrics further highlighted significant deficiencies in models’ logical consistency during reasoning. Our work also reveals that different models exhibit varying sensitivities to different categories, with non-visual commonsense understanding posing greater challenges than visual commonsense. This study aims to inspire further research to enhance the commonsense reasoning capabilities of artificial intelligence models.