Robot Perception: Sensors and Image Processing
摘要
This chapter explores how robots perceive their environment using various sensors and process visual data through CNNs and transformers for tasks like classification, segmentation, and object detection. It discusses the trade-offs between different CNN-based models and highlights the advantages of vision transformers (ViT) and detection transformers (DETR) in capturing global context. The chapter also covers scalability and emerging transformer-based methods beneficial for robotics applications.