<p>The development of Vision Language Models (VLMs) has made it possible to analyze the cultural dynamics of many cultural art forms from various nations. Robust VLMs have led to the emergence of novel applications of AI (artificial intelligence) and ML (machine learning), including visual question answering (VQA), image captioning, and general scene understanding. However, it is necessary to ensure that VLMs respect a variety of values, prevent ethical misalignments, and promote public trust by addressing accountability, and inclusivity in various cultural contexts. This work focuses on evaluating prominent VLMs in answering cultural queries of six traditional games of Assam, India. Prior works on deep learning-based cultural informatics over traditional games have not explored this dimension. This work proposes a novel evaluation paradigm which considers multidimensional assessment and pair-wise evaluation. It performs cross metric comparison and formulates three algorithms to analyze cultural comprehensibility of two VLMs. A visual question answering (VQA) dataset has been compiled which consists of 3267 image-question pairs for evaluating VLMs through our evaluation paradigm. Qwen2-VL-7B and Gemini 1.5 have been evaluated using three metrics (Cosine similarity and two varieties of LAVE accuracy). This work is going to help one in building an ethical question answering system on top of VLMs.</p>

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

Evaluating Vision Language Models (VLMs) across six traditional games: an evaluation paradigm to better assess culturally sensitive question answering systems

  • Rupam Nath,
  • Rupam Bhattacharyya

摘要

The development of Vision Language Models (VLMs) has made it possible to analyze the cultural dynamics of many cultural art forms from various nations. Robust VLMs have led to the emergence of novel applications of AI (artificial intelligence) and ML (machine learning), including visual question answering (VQA), image captioning, and general scene understanding. However, it is necessary to ensure that VLMs respect a variety of values, prevent ethical misalignments, and promote public trust by addressing accountability, and inclusivity in various cultural contexts. This work focuses on evaluating prominent VLMs in answering cultural queries of six traditional games of Assam, India. Prior works on deep learning-based cultural informatics over traditional games have not explored this dimension. This work proposes a novel evaluation paradigm which considers multidimensional assessment and pair-wise evaluation. It performs cross metric comparison and formulates three algorithms to analyze cultural comprehensibility of two VLMs. A visual question answering (VQA) dataset has been compiled which consists of 3267 image-question pairs for evaluating VLMs through our evaluation paradigm. Qwen2-VL-7B and Gemini 1.5 have been evaluated using three metrics (Cosine similarity and two varieties of LAVE accuracy). This work is going to help one in building an ethical question answering system on top of VLMs.