RSTAN: Residual Spatio-Temporal Attention Network for End-to-End Human Fall Detection
摘要
The occurrence of human fall is a significant threat to human health, especially among the elderly. Unlike standard action recognition, falls manifest a combination of static and dynamic attributes. They are highly sensitive to spatio-temporal motion, marked by sudden and transient occurrences. This paper proposes a novel spatio-temporal convolutional method for end-to-end human fall detection, named Residual Spatio-Temporal Attention Network (RSTAN). The network integrates a Spatial Channel Attention (SCA) module within the convolutional layers of the Residual 3D convolution to enhance feature refinement. selectively accentuates spatial and channel dimensions. In addition, to capture both the extensive spatio-temporal features and the short-range spatio-temporal characteristics of human falls, effectively distinguishing them from daily activities, we propose a Multi-interval Difference Aggregation (MDA) method. The MDA utilizes multiple time interval frame differences to extract motion features. Our proposed method’s superior performance is demonstrated through experiments on three publicly available fall detection datasets. Specifically, achieving 100% accuracy on the UR Fall Detection dataset.