Vision Transformers in Emotion and Demographic Analysis: A Review of Techniques and Applications
摘要
Multimodal emotion detection and age-gender estimation have gained increased attention due to their wide applications across various fields such as affective computing, human–computer interaction and personalized services. Few such techniques have been proposed in the past few years which capitalizes upon the advancements in deep learning and make these tasks much more reliable. One such technique is the Vision Transformer (ViT) model which has demonstrated state of art results on some existing computer vision tasks. Focussing specifically on Vision Transformer models, this review paper aims to explore and critically analyse existing approaches and methodology employed for multimodal emotion detection and age-gender estimation. Besides the more recent deep learning methods such as convolutional neural networks (CNNs) and recurrent neural networks (RNNs), we also examine traditional modes based on handcrafted features and shallow machine learning models, summarizing their pros, cons, as well as avenues for further research thus paving the way to future advances in this domain.