Pallet Detection and 3D Pose Estimation via Geometric Cues Learned from Synthetic Data
摘要
Vision-based object recognition is an important enabler for automating specific workflows in production and transportation scenarios. Locating and manipulating pallets as common functional objects represents a relevant robotic task in these domains. However, learning accurate neural models to estimate the location and pose of such objects is non-trivial. The main complexity stems from the underlying diversity of representing pallet objects: varied viewing conditions, frequent occlusions, self-occluded parts, diverse pallet materials jointly span a vast space of possible appearances. We present a solution which tackles the data diversity problem and the issue of occlusions. On one hand, we demonstrate how to rely on synthetic image pairs to compute geometry-encoding stereo disparity images, highly independent from appearance variations and exhibiting a small synthetic-to-real data domain gap. On the other hand, we introduce a novel point- and line-segment-based voting scheme, yielding a strong support for object presence in case of occlusions and up to distances of 8 m. We provide a quantitative evaluation of recognition accuracy for several network architectures using a manually fine-annotated multi-warehouse data-set. Based on the presented pallet recognition scheme, we also describe an automated forklift demonstrator, able to perform automated pallet pick-up and drop-off operations under diverse observation conditions.