Audio Situation Analysis by Multiple Signal Features on a Hybrid Parallel Neural Network
摘要
The human voice is one of the best possible ways to communicate and help the other person to understand better. According to Albert Mehrabian body language, our verbal communication consists of 38% vocal and 7% word only communication. Though there are 55% other ways of communication we need a prominent feature that is easy to detect and at the same time very important. The scientific community has focused on text-to-audio conversion of audio content like songs, debates, news, and political arguments. This paper proposes a hybrid parallel CNN and Attention (transformer) based architecture to represent the given audio into valuable information and classify the given situation. The model gets trained parallelly for the CNN and Transformer and then the output of the feature tensor is combined and fed into a combined model network passing through dense layers finally classifying the audio. The datasets used for the implementation purpose are EMOVO and Corpus.