FPSFT: Frame-Patch-Select Fusion Transformer for Action Recognition
摘要
Video action recognition plays a crucial role in numerous applications. However, current research encounters significant challenges, such as the difficulty in extracting features that maintain long-term connections within the video feature space and the inefficiency of attention computation. Therefore, we propose an innovative feature fusion method based on attention mechanism. This method employs the Transformer to create a Key Frame and Key Patch Selection module alongside a Small-Big Patch Transformer module. These components efficiently establish relationships between features with long-term connections in video data. By integrating these modules within the Transformer and combining them with Convolutional Neural Networks, our approach leverages the strengths of different frameworks. This integration significantly enhances the computational efficiency and accuracy. The experiments conducted on the generic video datasets demonstrate that our model surpasses the performance of most previous efforts at the same input scale.