This study examined the impact of AI-generated image descriptions on the classification of tourism-related images. It specifically compared two AI models, LLaVA and BLIP, and used BER Topic to cluster their generated descriptions. By analyzing two contrasting tourist destinations, Tokyo Disneyland and Kenrokuen, the research revealed that LLaVA is more effective for handling complex and diverse images, whereas BLIP excels with consistent imagery. These results show the significance of choosing an AI model that aligns with the dataset’s characteristics.

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

Improving Tourism Image Classification with AI-Generated Descriptions: A Comparative Analysis of LLaVA and BLIP Models

  • Suguru Tsujioka,
  • Kojiro Watanabe,
  • Akihiro Tsukamoto,
  • Muneyuki Natsume

摘要

This study examined the impact of AI-generated image descriptions on the classification of tourism-related images. It specifically compared two AI models, LLaVA and BLIP, and used BER Topic to cluster their generated descriptions. By analyzing two contrasting tourist destinations, Tokyo Disneyland and Kenrokuen, the research revealed that LLaVA is more effective for handling complex and diverse images, whereas BLIP excels with consistent imagery. These results show the significance of choosing an AI model that aligns with the dataset’s characteristics.