Object detection is a cornerstone of computer vision, yet the emergence of Vision–Language Models (VLMs) raises the question of whether these multimodal systems can rival specialized, fully supervised detectors. This paper presents a comparative study between a state-of-the-art VLM (Florence-2) and YOLOv8 on a vehicle-focused subset of COCO. We then benchmark both models using standard detection metrics. The results show that VLMs achieve competitive performance without triggers even though YOLOv8 maintains the highest overall precision. These results underline the potential of VLMs as a promising alternative when labelled data is scarce, and lay the groundwork for future research on rapid engineering, domain adaptation and evaluation in more challenging scenarios.

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

Assessing the Object-Detection Skills of Modern Vision Language Models

  • Paula Terleira Fernández,
  • Lucía Díez García,
  • André Sales Mendes,
  • Álvaro Lozano Murciego

摘要

Object detection is a cornerstone of computer vision, yet the emergence of Vision–Language Models (VLMs) raises the question of whether these multimodal systems can rival specialized, fully supervised detectors. This paper presents a comparative study between a state-of-the-art VLM (Florence-2) and YOLOv8 on a vehicle-focused subset of COCO. We then benchmark both models using standard detection metrics. The results show that VLMs achieve competitive performance without triggers even though YOLOv8 maintains the highest overall precision. These results underline the potential of VLMs as a promising alternative when labelled data is scarce, and lay the groundwork for future research on rapid engineering, domain adaptation and evaluation in more challenging scenarios.