<p>Facial Expression Recognition (FER) is a challenging problem in computer vision due to various factors, including illumination variability, occlusions, high intra-class similarity of expressions, and the complexity of accurately capturing subtle facial deformations. Traditional CNN-based models often suffer from spatial information loss and overfitting, especially when dealing with high-dimensional feature spaces. Furthermore, existing methods struggle to maintain performance consistency across real-time applications and diverse datasets. The proposed system introduces an AI-integrated, advanced FER framework based on a Hyperscale YOLOv5 Capsuled Convolutional Neural Network (YOLOv5c-CNN) combined with swarm intelligence optimization techniques. This novel methodology is structured into multiple stages to ensure high precision, robustness, and real-time performance in emotion detection from facial features. In the pre-processing stage, the input facial image undergoes normalization using a Non-Local-Median Collateral Filter (NL-MCF). This filter preserves edge integrity while minimizing noise and illumination variations, ensuring that essential facial features are retained for downstream processing. Subsequently, Adaptive Histogram Equalization is employed to enhance image contrast by computing the pixel intensity Cumulative Distribution Function (CDF), which further improves feature visibility in varied lighting conditions. The next stage involves Cross-Layer Slice Fragment Window Segmentation (CLSFWS). This segmentation technique isolates critical facial regions especially those prone to expression-induced deformations (e.g., eyes, mouth, brows) by detecting pitfall markings and slice-based layer variations. To tackle the high dimensionality of facial feature datasets and avoid overfitting, Adaptive Particle Swarm Optimization (APSO) is integrated as a feature selection mechanism. APSO dynamically selects the most relevant features by simulating intelligent swarm-based behavior, optimizing the feature space for improved classification accuracy and computational efficiency. The final classification and expression recognition are performed using the proposed Hyperscale YOLOv5 Capsuled Convolutional Neural Network (YOLOv5c-CNN). This architecture enhances the standard YOLOv5 model by integrating capsule layers, which preserve spatial hierarchies and relationships between facial features—offering superior recognition of subtle and complex expressions. The hyperscale feature refers to the ability to process large-scale facial datasets with minimal degradation in the model’s scalability and performance. The YOLOv5c-CNN model achieves the highest accuracy of 96.39% across all dataset sizes, demonstrating strong performance in facial expression classification.</p>

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

Ai integrated advanced facial expression recognition using hyperscale YOLOV5 capsuled convolution neural network with swarm intelligence technique

  • G. NandaGopal,
  • V. S. Arulmurugan,
  • M. Mohammadha Hussaini

摘要

Facial Expression Recognition (FER) is a challenging problem in computer vision due to various factors, including illumination variability, occlusions, high intra-class similarity of expressions, and the complexity of accurately capturing subtle facial deformations. Traditional CNN-based models often suffer from spatial information loss and overfitting, especially when dealing with high-dimensional feature spaces. Furthermore, existing methods struggle to maintain performance consistency across real-time applications and diverse datasets. The proposed system introduces an AI-integrated, advanced FER framework based on a Hyperscale YOLOv5 Capsuled Convolutional Neural Network (YOLOv5c-CNN) combined with swarm intelligence optimization techniques. This novel methodology is structured into multiple stages to ensure high precision, robustness, and real-time performance in emotion detection from facial features. In the pre-processing stage, the input facial image undergoes normalization using a Non-Local-Median Collateral Filter (NL-MCF). This filter preserves edge integrity while minimizing noise and illumination variations, ensuring that essential facial features are retained for downstream processing. Subsequently, Adaptive Histogram Equalization is employed to enhance image contrast by computing the pixel intensity Cumulative Distribution Function (CDF), which further improves feature visibility in varied lighting conditions. The next stage involves Cross-Layer Slice Fragment Window Segmentation (CLSFWS). This segmentation technique isolates critical facial regions especially those prone to expression-induced deformations (e.g., eyes, mouth, brows) by detecting pitfall markings and slice-based layer variations. To tackle the high dimensionality of facial feature datasets and avoid overfitting, Adaptive Particle Swarm Optimization (APSO) is integrated as a feature selection mechanism. APSO dynamically selects the most relevant features by simulating intelligent swarm-based behavior, optimizing the feature space for improved classification accuracy and computational efficiency. The final classification and expression recognition are performed using the proposed Hyperscale YOLOv5 Capsuled Convolutional Neural Network (YOLOv5c-CNN). This architecture enhances the standard YOLOv5 model by integrating capsule layers, which preserve spatial hierarchies and relationships between facial features—offering superior recognition of subtle and complex expressions. The hyperscale feature refers to the ability to process large-scale facial datasets with minimal degradation in the model’s scalability and performance. The YOLOv5c-CNN model achieves the highest accuracy of 96.39% across all dataset sizes, demonstrating strong performance in facial expression classification.