<p>In real-time video conferencing, ensuring high-quality, blur-free video while preserving facial micro-expressions and natural motion remains a challenge, particularly for AI-driven virtual meetings, telemedicine, interactive remote collaboration, and real-time facial recognition applications. To address this, we have introduced ViConNet, a novel network for real-time video deblurring and enhancement with facial micro-expression and pose estimation. Our framework integrates DINO-ViT for self-supervised feature extraction, ViTPose for precise head pose estimation, ensuring robust facial identity preservation and Restormer+, a transformer-based video deblurring mechanism that selectively refines motion-blurred regions and facial features, ensuring optimal temporal coherence and high perceptual quality. Recognizing the practical deployment challenges of deep models on edge devices, we further introduce TSPD (Temporal-Spatial Probe Distilled), a knowledge distillation framework that effectively transfers multi-scale spatial features, motion-aware embeddings, and facial micro-expression dynamics from a high-capacity teacher model to a lightweight student model. This distilled student model is specifically motivated by the need to achieve real-time performance on resource-constrained devices without sacrificing restoration quality. Experimental results on the 2MF<sup>2</sup>, YouTube Faces (YTF) and TalkingHead-1KH datasets demonstrate that our model outperforms state-of-the-art deblurring methods, achieving PSNR improvements of 3.21% and 2.76%, respectively. The student model maintains comparable PSNR (3.15% and 2.59% improvements) while achieving a 45% reduction in model size and 40% lower computational cost, enabling faster inference and broader applicability in real-world mobile and embedded environments.</p>

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

ViConNet: Temporal-Spatial Probe Distilled (TSPD) Network for Real-Time Video Deblurring and Enhancement with Facial Micro-expression and Pose Estimation for Real-Time Video Conferencing

  • Arti Ranjan,
  • M. Ravinder

摘要

In real-time video conferencing, ensuring high-quality, blur-free video while preserving facial micro-expressions and natural motion remains a challenge, particularly for AI-driven virtual meetings, telemedicine, interactive remote collaboration, and real-time facial recognition applications. To address this, we have introduced ViConNet, a novel network for real-time video deblurring and enhancement with facial micro-expression and pose estimation. Our framework integrates DINO-ViT for self-supervised feature extraction, ViTPose for precise head pose estimation, ensuring robust facial identity preservation and Restormer+, a transformer-based video deblurring mechanism that selectively refines motion-blurred regions and facial features, ensuring optimal temporal coherence and high perceptual quality. Recognizing the practical deployment challenges of deep models on edge devices, we further introduce TSPD (Temporal-Spatial Probe Distilled), a knowledge distillation framework that effectively transfers multi-scale spatial features, motion-aware embeddings, and facial micro-expression dynamics from a high-capacity teacher model to a lightweight student model. This distilled student model is specifically motivated by the need to achieve real-time performance on resource-constrained devices without sacrificing restoration quality. Experimental results on the 2MF2, YouTube Faces (YTF) and TalkingHead-1KH datasets demonstrate that our model outperforms state-of-the-art deblurring methods, achieving PSNR improvements of 3.21% and 2.76%, respectively. The student model maintains comparable PSNR (3.15% and 2.59% improvements) while achieving a 45% reduction in model size and 40% lower computational cost, enabling faster inference and broader applicability in real-world mobile and embedded environments.