The field of audio sentiment analysis, particularly through automatic speech recognition (ASR), is an emerging area of research. Unlike the well-established domain of text-based sentiment analysis, the detection of opinions or sentiments expressed by a speaker in natural audio remains relatively underexplored. This endeavor poses a unique set of challenges, making it a complex and intriguing problem to tackle. Extracting sentiment from audio sources involves deciphering the emotional nuances, tones, and expressions embedded in spoken language, adding a layer of difficulty compared to the more common text-based sentiment detection methods. Despite its challenges, this field holds promise for gaining deeper insights into human emotions and opinions conveyed through spoken words. This proposed work delves deep into the intricacies of decoding emotions from speech. The proposed work utilizes a diverse range of models, including LSTM networks, random forest, SVM, and CNN, to explore this complex landscape. Our study harnesses the power of the Toronto Emotional Speech Set (TESS) dataset, meticulously analyzing 2800 audio files to uncover the vast spectrum of emotions portrayed by two incredibly talented actresses. Each model in our ensemble showcases its unique strengths, with standout accuracy of 94% for LSTM. Our comprehensive evaluation goes beyond simple accuracy, encompassing precision, recall, and F1-score metrics, providing valuable insights into the distinctive abilities of each model. Through this proposed work, we not only showcase the remarkable performance of individual models but also emphasize the collective impact of employing a range of approaches to achieve a more nuanced understanding of emotion decoding in speech.

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

Decoding Emotions Through Speech Using LSTM

  • Tejas Nadagadalli,
  • Khushi Patil,
  • Supreet Palankar,
  • Pragati Bhat

摘要

The field of audio sentiment analysis, particularly through automatic speech recognition (ASR), is an emerging area of research. Unlike the well-established domain of text-based sentiment analysis, the detection of opinions or sentiments expressed by a speaker in natural audio remains relatively underexplored. This endeavor poses a unique set of challenges, making it a complex and intriguing problem to tackle. Extracting sentiment from audio sources involves deciphering the emotional nuances, tones, and expressions embedded in spoken language, adding a layer of difficulty compared to the more common text-based sentiment detection methods. Despite its challenges, this field holds promise for gaining deeper insights into human emotions and opinions conveyed through spoken words. This proposed work delves deep into the intricacies of decoding emotions from speech. The proposed work utilizes a diverse range of models, including LSTM networks, random forest, SVM, and CNN, to explore this complex landscape. Our study harnesses the power of the Toronto Emotional Speech Set (TESS) dataset, meticulously analyzing 2800 audio files to uncover the vast spectrum of emotions portrayed by two incredibly talented actresses. Each model in our ensemble showcases its unique strengths, with standout accuracy of 94% for LSTM. Our comprehensive evaluation goes beyond simple accuracy, encompassing precision, recall, and F1-score metrics, providing valuable insights into the distinctive abilities of each model. Through this proposed work, we not only showcase the remarkable performance of individual models but also emphasize the collective impact of employing a range of approaches to achieve a more nuanced understanding of emotion decoding in speech.