Vision transformer and graph neural networks based SuryaNet for yoga pose quality assessment
摘要
There has been significant growth in the number of people practicing Yoga, and as a result there is a growing need for advanced recognition systems and systems to measure how well a user performs their yoga postures. Most current recognition systems have demonstrated high levels of success in classifying yoga postures but none of these systems assess the quality of a user’s execution of a posture. This paper therefore introduces SuryaNet, a multi-task model with four technical contributions: (1) A Vision Transformer backbone that captures global spatial dependency relationships; (2) An Adaptive Spatial-Temporal Graph Convolutional Network (ST-GCNN) that models the skeletal topology; (3) Bidirectional Cross-Modal Transformer Fusion with learned gating for adaptive modality weighting; (4) Quality-Aware Contrastive Learning with Uncertainty-Weighted Multi-Task Optimization that jointly trains pose classification and continuous quality scoring. The results demonstrate that SuryaNet obtains 99.36% classification accuracy and 0.9547 Pearson correlation for quality prediction, which establishes a new state-of-the-art in terms of quality assessment, while it also provides users with quality feedback.