<p>Facial emotion recognition (FER) is a prominent area of research in computer vision, yet it remains difficult to detect a person's expressions in a complex environment. Another vital challenge is how to reduce computations and decrease the training process while retaining good accuracy with limited hardware resources. In this paper, on the basis of the facial landmark approaches, a cross connection with depthwise separable convolutions (DSC), pixel shuffle, and an anti-aliasing model is proposed, viz. DC-PAD. Davis’s Library (Dlib) is used first for extracting facial landmarks to decrease redundancy in facial data. Following that, these landmarks will be fed as input to the DC-PAD model. DC-PAD utilizes depthwise and pointwise convolutions; the anti-aliasing technique is applied to remove the intermediate artifacts resulting from downsampling operations, while the Pixel Shuffle operation has been deployed to blend different-sized feature maps. The proposed model has been carried out and evaluated on three commonly used FER datasets to detect eight different emotions from image data, including contempt, happiness, neutrality, disgust, sadness, fear, surprise, and anger. These results will then be contrasted with another approach (DC-PAC), which employs the typical convolutional layers instead of the DSC layers. Furthermore, to measure the reliability and efficiency of the proposed approach, numerous performance metrics involving accuracy, specificity, sensitivity, Jaccard coefficient, training time, and the overall number of parameters have been evaluated after applying each operation. On the Real-world Affective Face dataset, one of the most challenging FER datasets, the proposed DC-PAD model obtains an overall accuracy of 83.89%, 99.03% on the Japanese female facial expressions (JAFFE), and 99.34% on the Extended Cohn–Kanade (CK+) datasets. The experimental results demonstrate that the DSC-based approach (DC-PAD) significantly outperforms in terms of reliability and precision while decreasing the dimensionality of the input data, reducing the training time, and minimizing the total model parameters.</p>

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

Supercomputing-efficient facial emotion recognition: reducing parameter complexity with pixel shuffle and anti-aliasing in depthwise separable models

  • Reham A. Elsheikh,
  • M. A. Mohamed,
  • Ahmed Mohamed Abou-Taleb,
  • Mohamed Maher Ata

摘要

Facial emotion recognition (FER) is a prominent area of research in computer vision, yet it remains difficult to detect a person's expressions in a complex environment. Another vital challenge is how to reduce computations and decrease the training process while retaining good accuracy with limited hardware resources. In this paper, on the basis of the facial landmark approaches, a cross connection with depthwise separable convolutions (DSC), pixel shuffle, and an anti-aliasing model is proposed, viz. DC-PAD. Davis’s Library (Dlib) is used first for extracting facial landmarks to decrease redundancy in facial data. Following that, these landmarks will be fed as input to the DC-PAD model. DC-PAD utilizes depthwise and pointwise convolutions; the anti-aliasing technique is applied to remove the intermediate artifacts resulting from downsampling operations, while the Pixel Shuffle operation has been deployed to blend different-sized feature maps. The proposed model has been carried out and evaluated on three commonly used FER datasets to detect eight different emotions from image data, including contempt, happiness, neutrality, disgust, sadness, fear, surprise, and anger. These results will then be contrasted with another approach (DC-PAC), which employs the typical convolutional layers instead of the DSC layers. Furthermore, to measure the reliability and efficiency of the proposed approach, numerous performance metrics involving accuracy, specificity, sensitivity, Jaccard coefficient, training time, and the overall number of parameters have been evaluated after applying each operation. On the Real-world Affective Face dataset, one of the most challenging FER datasets, the proposed DC-PAD model obtains an overall accuracy of 83.89%, 99.03% on the Japanese female facial expressions (JAFFE), and 99.34% on the Extended Cohn–Kanade (CK+) datasets. The experimental results demonstrate that the DSC-based approach (DC-PAD) significantly outperforms in terms of reliability and precision while decreasing the dimensionality of the input data, reducing the training time, and minimizing the total model parameters.