Assessing the Object-Detection Skills of Modern Vision Language Models
摘要
Object detection is a cornerstone of computer vision, yet the emergence of Vision–Language Models (VLMs) raises the question of whether these multimodal systems can rival specialized, fully supervised detectors. This paper presents a comparative study between a state-of-the-art VLM (Florence-2) and YOLOv8 on a vehicle-focused subset of COCO. We then benchmark both models using standard detection metrics. The results show that VLMs achieve competitive performance without triggers even though YOLOv8 maintains the highest overall precision. These results underline the potential of VLMs as a promising alternative when labelled data is scarce, and lay the groundwork for future research on rapid engineering, domain adaptation and evaluation in more challenging scenarios.