Automated Contextual Tagging in Points-of-Interest Using Bag-of-Objects
摘要
Object detection is a fundamental task in computer vision, with its applications ranging from autonomous driving to scene recognition. In the domain of Points-of-Interest (POI), object recognition can aid in the analysis of complex urban environments that contain various types of infrastructure, people, and activities. This paper addresses this challenge by proposing a Bag-of-Objects (BoO) methodology for POI scene contextual tagging. The suggested approach utilizes a transformer-based model for object detection, followed by a threshold to assign tags to POI scenes. These tags are then used to train a novel multimodal model, called WORLD (Weight Optimization for Representation and Labeling Descriptions), which is capable of classifying and contextually tagging POI scenes based on their visual features.