Enhancing Mobile Manipulation in Home Environments: A Case Study from the NeurIPS 2023 HomeRobot Challenge
摘要
We present enhancements to the reinforcement learning (RL) approach used in the NeurIPS 2023 HomeRobot: Open Vocabulary Mobile Manipulation (OVMM) Challenge, focusing on augmenting the baseline model with advanced semantic segmentation and skill policy modifications. More specifically, we introduce refined semantic segmentation model (integrating the YOLOv8 and MobileSAM), improved place skill policy and a high-level heuristic strategy, which collectively advance the overall success rate from 0.8 to 5.2 (+550% relative) and the partial success rate from 9.7 to 25.8 (+165% relative) on the Test Standard split of the challenge dataset (ranked 2nd on the public leaderbord). These enhancements enabled our agent to achieve 3rd place in both the simulated and real-world stages of the competition. This paper details the strategies employed, discusses the insights gained, particularly in semantic segmentation and skill-specific training, and outlines potential avenues for future enhancements in embodied AI systems within open-vocabulary contexts.