<p>Facial emotion detection has been receiving widespread interest due to its broad applications in various fields. Nevertheless, unifying disparate data modalities, like visual and text data, is a complicated procedure. This research introduces a Voting-based Optimized Deep Ensemble (VODE) Model designed to solve these challenges by leveraging base learners and meta learners to enhance facial emotion recognition accuracy. This model incorporates advanced feature extraction modules, involving Scale-Invariant Feature Transform (SIFT) for visual data and FastText for textual data, to effectively capture key emotional cues. A multi-head transformer model is developed to efficiently fuse these features, capable of capturing intricate relationships between different modalities. Additionally, it utilizes an ensemble learning model by combining predictions from multiple deep networks, such as VGG19, DenseNet121, ResNetv2, InceptionResNet, and MobileNet, with fine-tuned support vector machine and gradient boosting algorithms for emotion recognition These ensemble predictions are then optimized through a majority voting mechanism, further enhancing the model's performance. The experimental outcomes establish significant improvements in emotion recognition, also validating the efficiency of the introduced model. The approach attained 98% of accuracy, 98% of precision, and 98.5% of recall. These metrics underscore the higher functionality of the introduced approaches when compared to existing techniques.</p>

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

Voting based optimized deep ensemble model: an effective visual and textual fusion for recognition of face emotions

  • Dipti Pandit,
  • Sangeeta Jadhav

摘要

Facial emotion detection has been receiving widespread interest due to its broad applications in various fields. Nevertheless, unifying disparate data modalities, like visual and text data, is a complicated procedure. This research introduces a Voting-based Optimized Deep Ensemble (VODE) Model designed to solve these challenges by leveraging base learners and meta learners to enhance facial emotion recognition accuracy. This model incorporates advanced feature extraction modules, involving Scale-Invariant Feature Transform (SIFT) for visual data and FastText for textual data, to effectively capture key emotional cues. A multi-head transformer model is developed to efficiently fuse these features, capable of capturing intricate relationships between different modalities. Additionally, it utilizes an ensemble learning model by combining predictions from multiple deep networks, such as VGG19, DenseNet121, ResNetv2, InceptionResNet, and MobileNet, with fine-tuned support vector machine and gradient boosting algorithms for emotion recognition These ensemble predictions are then optimized through a majority voting mechanism, further enhancing the model's performance. The experimental outcomes establish significant improvements in emotion recognition, also validating the efficiency of the introduced model. The approach attained 98% of accuracy, 98% of precision, and 98.5% of recall. These metrics underscore the higher functionality of the introduced approaches when compared to existing techniques.