Reinforcement Learning-Driven Adaptive Coaching using 3D Vision Transformers for Personalized Yoga Pose Correction
摘要
Artificial intelligence-assisted wellness systems have significantly improved digital fitness training, yet most existing yoga posture recognition frameworks primarily focus on pose classification and provide only static corrective feedback. Such systems often fail to adapt to individual body flexibility, posture progression, and user-specific biomechanical variations during practice. To address these limitations, this paper proposes a Reinforcement Learning-Driven Adaptive Coaching framework using 3D Vision Transformers for personalized yoga pose correction. The proposed system utilizes monocular RGB video input to estimate three-dimensional skeletal joint coordinates through a robust 3D human pose reconstruction pipeline. A Spatial–Temporal Vision Transformer is employed to model complex spatial relationships between joints and temporal dependencies across yoga movements. Unlike conventional posture recognition methods, the proposed framework integrates a reinforcement learning agent that continuously evaluates user posture quality and generates adaptive corrective guidance based on biomechanical feedback. A biomechanical rewartrd mechanism is designed to assess posture alignment, joint-angle accuracy, symmetry, and balance stability during yoga practice. These reward signals guide the reinforcement learning policy to optimize personalized coaching strategies for different users and varying posture difficulties. The system dynamically adjusts correction intensity and feedback according to the practitioner’s performance, thereby enabling progressive and individualized training. To improve robustness under real-world conditions, the framework incorporates temporal smoothing, viewpoint normalization, and occlusion-aware skeletal modeling. Experimental evaluation is conducted using Yoga-82, Human3.6 M, and a custom multi-subject yoga posture dataset containing variations in camera viewpoints, illumination conditions, and body structures. The proposed framework demonstrates superior posture recognition accuracy, reduced joint-angle deviation, and improved corrective feedback efficiency when compared with conventional CNN, LSTM, and graph-based baselines. Furthermore, the reinforcement learning module significantly enhances long-term posture correction performance by learning adaptive coaching policies from continuous user interaction. The proposed framework offers an intelligent, scalable, and real-time solution for personalized yoga training, rehabilitation support, and next-generation digital healthcare applications. The proposed method achieved 97.4% pose recognition accuracy, 24.9 MPJPE, 96.8% PCK, and 29 FPS real-time processing, outperforming existing CNN, CNN-LSTM, HRNet, and Vision Transformer-based approaches in the revised version of the manuscript.