Though existing 3D human pose estimation approaches achieve significant performance improvements, they struggle under occlusions. To mitigate this challenge, we propose OCCL-Former, a novel TransFormer-based OCCLusion method that implicitly learns the dependencies and relationships among body parts to accurately infer occluded body regions exploiting visible information. Firstly, OCCL-Former estimates the interrelationships between body parts and joints during training using augmented inputs. By leveraging augmented data to extract global contextual information and dependencies, OCCL-Former robustly handles occlusions. During training, we simulate occlusions by introducing noise to obscure specific body parts while providing full-body visuals as targets, thereby enabling the model to learn comprehensive inter-joint dependencies. The proposed architecture employs two transformers network, where one is for RGB inputs and another for partially segmented inputs and aggregating their outputs to produce precise pose and shape estimations. Experimental results demonstrate that OCCL-Former surpasses existing state-of-the-art methods, delivering superior accuracy and robustness in 3D human pose estimation under both standard and occluded conditions on standard benchmarks. The proposed OCCL-Former achieves 9.6, 9.7, and 8.4 lesser MPJPE, PA-MPJPE, and PVE error, respectively than the recent state-of-the-art approach on the popular benchmark standard 3DPW dataset.

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

OCCL-Former: Data Augmentation Driven Occlusion-Aware Inter-Body Parts Relationship Learning for 3D Pose Estimation

  • Md. Imtiaz Hossain,
  • Sharmen Akhter,
  • Sungjun Yang,
  • Eui-Nam Huh

摘要

Though existing 3D human pose estimation approaches achieve significant performance improvements, they struggle under occlusions. To mitigate this challenge, we propose OCCL-Former, a novel TransFormer-based OCCLusion method that implicitly learns the dependencies and relationships among body parts to accurately infer occluded body regions exploiting visible information. Firstly, OCCL-Former estimates the interrelationships between body parts and joints during training using augmented inputs. By leveraging augmented data to extract global contextual information and dependencies, OCCL-Former robustly handles occlusions. During training, we simulate occlusions by introducing noise to obscure specific body parts while providing full-body visuals as targets, thereby enabling the model to learn comprehensive inter-joint dependencies. The proposed architecture employs two transformers network, where one is for RGB inputs and another for partially segmented inputs and aggregating their outputs to produce precise pose and shape estimations. Experimental results demonstrate that OCCL-Former surpasses existing state-of-the-art methods, delivering superior accuracy and robustness in 3D human pose estimation under both standard and occluded conditions on standard benchmarks. The proposed OCCL-Former achieves 9.6, 9.7, and 8.4 lesser MPJPE, PA-MPJPE, and PVE error, respectively than the recent state-of-the-art approach on the popular benchmark standard 3DPW dataset.