Scene-aware contrastive regression for multi-person action quality assessment
摘要
Multi-person Action Quality Assessment (AQA) is crucial for evaluating group-based physical activities, particularly in educational and health-related contexts. However, current AQA research predominantly focuses on single-person sports scenarios, lacking the capability to evaluate multiple participants performing long-duration synchronized exercises. To address this gap, we introduce MELO, the first long-duration dataset specifically designed for multi-person AQA as individuals, featuring 192 long-duration videos (averaging 283 seconds) of children performing the Eight Section Brocade exercise, with comprehensive annotations for individual identities and action quality scores. We also propose MReC, a novel framework that addresses the unique challenges of multi-person AQA through three key components: (1) a temporal-spatial cross-shaped window self-attention mechanism (T-CSWin) for capturing local contextual relationships in group performances, (2) a selective fusion module (SFM) that integrates both video motion data and pose information for a more comprehensive assessment, and (3) a temporal-aware regression (TAR) module for effective long-duration score generation. Additionally, we introduce a scene-aware pairwise contrastive regression (SPCR) approach that enhances model performance by analyzing subtle motion differences between participants within the same scene during training. Extensive experiments demonstrate that MReC effectively handles the complexities of group-based evaluation, achieving superior performance on the MELO dataset and establishing new benchmarks for multi-person AQA in real-world educational settings.