An Appearance-based VisionTransformer Network for Enhanced Gaze Estimation
摘要
Appearance-based gaze estimation, which analyses the visual characteristics of the face, provides crucial hints for comprehending human intentions, especially in the domains of virtual reality and human-computer interaction. It also reveals where an individual focuses its visual attention. We present, in this article, a new gaze estimation method, called "Hybrid ViTNet" which combines two powerful models, ResNet50 and the Transformer, to improve feature representation in eye images. We have used ResNet50 for feature extraction, taking advantage of its ability to capture complex local details. By integrating the Transformer with the Multi-Head-Self-Attention layer, our model takes long-distance relationships into account, thereby improving overall context understanding. These two components are then reinforced by a feedforward network. Our proposed approach focuses on accurate classification of eye images for gaze direction estimation. Evaluated on the MPIIGaze, Columbia Gaze and EyeDiap datasets, it shows promising performance with average angular errors of