Fog computing-driven logistics: leveraging few-shot learning and foundational computer vision models
摘要
This work introduces an advanced computer vision system designed for real-time auto-identification (Auto-ID) of loads in industrial and logistics environments, optimized through a fog computing architecture. The system monitors multiple docks in logistics warehouses, identifying pallet loads in real-time and providing feedback for efficient truck loading to prevent errors. It enhances traditional Auto-ID by incorporating pallet class classification, detection, and size estimation, utilizing foundational models, such as DINOv2, MobileNetV3, SAM, Depth Anything V2, and Depth Pro, alongside few-shot learning to generate efficient training datasets with minimal labeling effort. We propose three key innovations: (1) an embedding analysis approach for precise load classification; (2) DINOv2-based visual feature detection solutions for bounding box estimation, and (3) a depth-guided segmentation strategy for improved load isolation and measurement accuracy. Experimental results on a curated industrial dataset compiled by us demonstrate high mean average precision, with optimized trade-offs between latency and accuracy for fog computing constraints. This solutions reduces cloud dependency, supports rapid online updates, and ensures robust performance in dynamic logistics settings, making it a scalable solution for automated image labeling and load identification in industrial applications.