The development of this paper aims to create an emotional recognition framework for real-time occupational health analysis, applying a multimodal method based on Deep Learning. Initially, an exhaustive research process is undertaken on emotional recognition methods and techniques, analyzing their impact within the occupational health context. Subsequently, definitions of recognition are established, exploring research sources and successful cases related to similar processes. For the framework development, emotional recognition techniques based on facial expressions (FER), speech-emotional recognition (SER), and predefined datasets such as MESD, RAVDESS, TESS, and SAVEE are employed. These datasets enable the identification of physiological, gestural, non-verbal, and vocal signals through emotional extraction. These data sets serve as the foundation for training and validating a multi-layered convolutional neural network model. Using standard evaluation metrics like the confusion matrix and “Hold-out” validation techniques, a model is developed with 92% accuracy across the seven universal emotions. Machine Learning techniques are employed based on specific needs, and predictive models are implemented through a user-friendly web application.

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

Emotion Recognition Framework for Real-Time Occupational Health Analysis, Applying a Multimodal Method Based on Deep Learning

  • Andrés Noboa,
  • Elizabeth Morejón,
  • Freddy Tapia

摘要

The development of this paper aims to create an emotional recognition framework for real-time occupational health analysis, applying a multimodal method based on Deep Learning. Initially, an exhaustive research process is undertaken on emotional recognition methods and techniques, analyzing their impact within the occupational health context. Subsequently, definitions of recognition are established, exploring research sources and successful cases related to similar processes. For the framework development, emotional recognition techniques based on facial expressions (FER), speech-emotional recognition (SER), and predefined datasets such as MESD, RAVDESS, TESS, and SAVEE are employed. These datasets enable the identification of physiological, gestural, non-verbal, and vocal signals through emotional extraction. These data sets serve as the foundation for training and validating a multi-layered convolutional neural network model. Using standard evaluation metrics like the confusion matrix and “Hold-out” validation techniques, a model is developed with 92% accuracy across the seven universal emotions. Machine Learning techniques are employed based on specific needs, and predictive models are implemented through a user-friendly web application.