A Novel Fatigue Detection Method Based on Video Transformer
摘要
This paper presents a novel fatigue detection method based on a video Transformer, specifically designed for detecting driver fatigue. Unlike traditional Transformers, the proposed method introduces innovations in feature extraction, attention mechanism design, and end-to-end architecture. Notably, it enhances the model’s understanding of overall image layout and fine-grained details by introducing conditional positional encoding, thereby improving the accuracy of fatigue feature extraction. Additionally, the method employs a factorized dot-product attention mechanism, effectively reducing computational complexity while ensuring robust temporal feature extraction. Furthermore, a feature scaling module is incorporated to more comprehensively perceive facial movements. The end-to-end deep learning architecture is well-suited to capturing the dynamic information relevant to fatigue detection, enhancing both efficiency and accuracy. The model’s generalization ability and stability have been validated through extensive cross-validation on various public datasets. In summary, the proposed method significantly enhances the accuracy and efficiency of fatigue detection, providing reliable technical support for the field.