<p>To address the challenges of age and gender recognition in uncontrolled scenarios with facial absence or severe occlusion, this paper proposes a Spatial Correlation Guided Cross Scale Feature Fusion Network (SCGNet). The proposed method specifically tackles the limitations of existing approaches that heavily rely on facial features, which become unreliable under partial/complete occlusion scenarios. The method integrates multi-granularity semantic features through a Cross-Scale Combination (CSC) module, enhances local detail representation using a Local Feature Guided Fusion (LFGF) module, and designs a Spatial Correlation Composition Analysis (SCCA) module based on Getis-Ord Gi* statistics for feature reorganization, effectively resolving interference from non-informative regions. The SCCA module introduces a novel bipartite grouping mechanism that leverages hotspot detection to preserve discriminative body features when facial cues are unavailable. Comprehensive experiments demonstrate that SCGNet achieves state-of-the-art performance with minimum Mean Absolute Error (MAE) 4.01% for age estimation on IMDB-Clean (2.9% improvement over VOLO-D1) and highest gender classification accuracy on IMDB-Clean, UTKFace, and Lagenda datasets, showing improvements in cross-scene adaptability compared to VOLO and MiVOLO models respectively. Notably, the method maintains gender discrimination accuracy under complete facial occlusion scenarios, validating the effectiveness of spatial correlation modeling for non-facial feature reasoning, maintaining 97.32% gender accuracy even with complete facial occlusion on Lagenda dataset. The proposed architecture shows 73.30% CS@5 for age prediction in cross-domain testing, demonstrating superior cross-scene adaptability compared to VOLO (69.72%) and MiVOLO (71.27%). Ablation studies confirm the individual contributions of CSC, LFGF, and SCCA modules. This research provides new insights for robust identity analysis in human-computer interaction and intelligent security applications.</p>

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

Spatial correlation guided cross scale feature fusion for age and gender estimation

  • Shiyi Jiang,
  • Qing Ji,
  • Hukui Shi,
  • Che Chen,
  • Yang Xu

摘要

To address the challenges of age and gender recognition in uncontrolled scenarios with facial absence or severe occlusion, this paper proposes a Spatial Correlation Guided Cross Scale Feature Fusion Network (SCGNet). The proposed method specifically tackles the limitations of existing approaches that heavily rely on facial features, which become unreliable under partial/complete occlusion scenarios. The method integrates multi-granularity semantic features through a Cross-Scale Combination (CSC) module, enhances local detail representation using a Local Feature Guided Fusion (LFGF) module, and designs a Spatial Correlation Composition Analysis (SCCA) module based on Getis-Ord Gi* statistics for feature reorganization, effectively resolving interference from non-informative regions. The SCCA module introduces a novel bipartite grouping mechanism that leverages hotspot detection to preserve discriminative body features when facial cues are unavailable. Comprehensive experiments demonstrate that SCGNet achieves state-of-the-art performance with minimum Mean Absolute Error (MAE) 4.01% for age estimation on IMDB-Clean (2.9% improvement over VOLO-D1) and highest gender classification accuracy on IMDB-Clean, UTKFace, and Lagenda datasets, showing improvements in cross-scene adaptability compared to VOLO and MiVOLO models respectively. Notably, the method maintains gender discrimination accuracy under complete facial occlusion scenarios, validating the effectiveness of spatial correlation modeling for non-facial feature reasoning, maintaining 97.32% gender accuracy even with complete facial occlusion on Lagenda dataset. The proposed architecture shows 73.30% CS@5 for age prediction in cross-domain testing, demonstrating superior cross-scene adaptability compared to VOLO (69.72%) and MiVOLO (71.27%). Ablation studies confirm the individual contributions of CSC, LFGF, and SCCA modules. This research provides new insights for robust identity analysis in human-computer interaction and intelligent security applications.