Inferring human personality traits from brief video clips presents substantial complexity due to the intertwined nature of visual, audio, and linguistic signals. To address this, we introduce MoRE (Multisource Representation Encoder), an innovative architecture designed to efficiently process and integrate diverse sensory information for personality profiling. On the visual side, we establish a face-centric relational map and employ a Dual-Stream Structural-Appearance Network, which merges topology-aware message propagation modules with conventional spatial feature extractors. This setup allows us to simultaneously capture geometric configurations and fine-grained visual textures. Additionally, holistic cues and identity-related embeddings are derived using fixed-feature extractors based on lightweight residual and facial recognition networks. To represent temporal patterns, sequential frame embeddings are handled using a bidirectional gated recurrence module enhanced by dynamic temporal focus layers. Parallel to this, sound-based features are produced through a pre-trained perceptual audio encoder, while language context is encoded using a cross-lingual semantic transformer. For the fusion of all modalities, we devise an adaptive channel selection mechanism that emphasizes informative streams before feeding the unified representation into a regression module built upon stacked nonlinear projection layers. Experimental evaluations on diverse personality benchmarks confirm that MoRE achieves superior predictive performance and demonstrates strong adaptability across datasets.

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

MoRE: Structured Multisignal Encoding for Human Disposition Recognition from Short Media Clips

  • Huanzhen Zhang,
  • Yuhang Li,
  • Chengwei Ye,
  • Yufei Lin,
  • Bohan Hu

摘要

Inferring human personality traits from brief video clips presents substantial complexity due to the intertwined nature of visual, audio, and linguistic signals. To address this, we introduce MoRE (Multisource Representation Encoder), an innovative architecture designed to efficiently process and integrate diverse sensory information for personality profiling. On the visual side, we establish a face-centric relational map and employ a Dual-Stream Structural-Appearance Network, which merges topology-aware message propagation modules with conventional spatial feature extractors. This setup allows us to simultaneously capture geometric configurations and fine-grained visual textures. Additionally, holistic cues and identity-related embeddings are derived using fixed-feature extractors based on lightweight residual and facial recognition networks. To represent temporal patterns, sequential frame embeddings are handled using a bidirectional gated recurrence module enhanced by dynamic temporal focus layers. Parallel to this, sound-based features are produced through a pre-trained perceptual audio encoder, while language context is encoded using a cross-lingual semantic transformer. For the fusion of all modalities, we devise an adaptive channel selection mechanism that emphasizes informative streams before feeding the unified representation into a regression module built upon stacked nonlinear projection layers. Experimental evaluations on diverse personality benchmarks confirm that MoRE achieves superior predictive performance and demonstrates strong adaptability across datasets.