<p>Despite recent advances in deep vision models, estimating chronological age from unconstrained facial images remains challenging: age cues are sparse and localized, whereas real-world images vary widely in pose, illumination, and quality. Although deep convolutional networks have long been the workhorse for this task, recent vision transformers improve global reasoning but tend to diffuse attention over many low–value regions. To address this gap, we propose a new processing methodology and develop a novel inference engine, hereafter referred to as <Emphasis FontCategory="NonProportional">REFINE</Emphasis> (<i>Region- and Feature-INformed iNference Engine</i>), a content-adaptive module that selectively extracts and aggregates a small set of age-informative patches and fuses them with global semantics. Concretely, we couple <Emphasis FontCategory="NonProportional">REFINE</Emphasis> with the compact hybrid backbone <Emphasis FontCategory="NonProportional">CoAtNet-0</Emphasis>. At its intermediate high-resolution stage, which predicts dense saliency and scale maps, the top–<i>K</i> locations are selected via differentiable region-of-interest (ROI) sampling. The resulting tokens are aggregated by a lightweight transformer, and a cross-attention fusion block conditions the global context from <Emphasis FontCategory="NonProportional">CoAtNet-0</Emphasis> on these encoded patch tokens. The fused representation drives complementary heads for Gaussian regression, label-distribution prediction, and ordinal classification, yielding calibrated and robust estimates. On <Emphasis FontCategory="NonProportional">MORPH&#xa0;II</Emphasis>, <Emphasis FontCategory="NonProportional">UTKFace</Emphasis>, and <Emphasis FontCategory="NonProportional">Adience</Emphasis>, selective fusion of regional evidence with global context delivers consistent gains over a strong <Emphasis FontCategory="NonProportional">CoAtNet-0</Emphasis> baseline while preserving computational efficiency.</p>

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

Region- and feature-informed inference engine (REFINE) with CoAtNet-0 for facial age estimation

  • Cilia Azni,
  • Mohammed Khammari

摘要

Despite recent advances in deep vision models, estimating chronological age from unconstrained facial images remains challenging: age cues are sparse and localized, whereas real-world images vary widely in pose, illumination, and quality. Although deep convolutional networks have long been the workhorse for this task, recent vision transformers improve global reasoning but tend to diffuse attention over many low–value regions. To address this gap, we propose a new processing methodology and develop a novel inference engine, hereafter referred to as REFINE (Region- and Feature-INformed iNference Engine), a content-adaptive module that selectively extracts and aggregates a small set of age-informative patches and fuses them with global semantics. Concretely, we couple REFINE with the compact hybrid backbone CoAtNet-0. At its intermediate high-resolution stage, which predicts dense saliency and scale maps, the top–K locations are selected via differentiable region-of-interest (ROI) sampling. The resulting tokens are aggregated by a lightweight transformer, and a cross-attention fusion block conditions the global context from CoAtNet-0 on these encoded patch tokens. The fused representation drives complementary heads for Gaussian regression, label-distribution prediction, and ordinal classification, yielding calibrated and robust estimates. On MORPH II, UTKFace, and Adience, selective fusion of regional evidence with global context delivers consistent gains over a strong CoAtNet-0 baseline while preserving computational efficiency.