A Review of Human Mesh Reconstruction: Beyond 2D Video Object Segmentation
摘要
Video object segmentation aims to extract 2D object masks by segmenting video frames into multiple objects, which is crucial in various practical applications such as medical imaging, etc.. However, traditional video object segmentation methods produce 2D masks, which are not suitable for 3D scenarios where depth information is essential, such as in robotic grasping, virtual reality, and autonomous driving, etc.. In this paper, we present a comprehensive review of 3D human mesh reconstruction (HMR) as an extension beyond 2D video object segmentation. We begin by reviewing the mainstream video object segmentation methods, then transition from 2D video object segmentation to 3D HMR. We further categorize recent HMR methods based on key characteristics that define this research field, including the type of model input and the use of statistical models. Finally, we provide detailed information on HMR datasets and evaluation metrics.