<p>Livestream e-commerce now exceeds USD 500 billion in annual online retail sales. Nevertheless, a large proportion of viewers watch livestreams without sound, reducing engagement and lowering purchase conversion rates. In this work, we explore vision-based content understanding through lip reading and visual speech recognition (VSR) for automatic subtitle generation to enhance livestream clarity. We present LiveVSR-Net, a novel architecture combining 3D convolutional neural networks with transformer-based temporal modeling and hybrid CTC-attention decoding. Domain-adapted enhancements include e-commerce-specific vocabulary expansion, fine-tuned language model integration, and multi-speaker handling techniques. We conduct comprehensive ablation studies to quantify the contribution of each architectural component, evaluate on standard benchmarks (LRS2, LRS3, GRID), and provide real-world robustness analysis with deployability guidelines for livestreaming. Across three standard VSR benchmarks, the proposed LiveVSR-Net framework achieves competitive performance: 23.5% WER on LRS2, 25.1% WER on LRS3, and 0.9% on GRID. By combining 3D convolutional visual encoders, Conformer-based temporal modeling, hybrid CTC-attention decoding, and domain-adapted language models, the method achieves real-time recognition suitable for deployment in production environments.</p>

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

E-commerce video & content understanding via vision: a comprehensive study on lip reading and visual speech recognition for enhanced shopping experience

  • Hau Nguyen Trung,
  • Vinh Truong Hoang,
  • Nghia Dinh,
  • Luu Quang Phuong,
  • Ha Duong Thi Hong,
  • Bay Nguyen Van,
  • Thien Ho Huong

摘要

Livestream e-commerce now exceeds USD 500 billion in annual online retail sales. Nevertheless, a large proportion of viewers watch livestreams without sound, reducing engagement and lowering purchase conversion rates. In this work, we explore vision-based content understanding through lip reading and visual speech recognition (VSR) for automatic subtitle generation to enhance livestream clarity. We present LiveVSR-Net, a novel architecture combining 3D convolutional neural networks with transformer-based temporal modeling and hybrid CTC-attention decoding. Domain-adapted enhancements include e-commerce-specific vocabulary expansion, fine-tuned language model integration, and multi-speaker handling techniques. We conduct comprehensive ablation studies to quantify the contribution of each architectural component, evaluate on standard benchmarks (LRS2, LRS3, GRID), and provide real-world robustness analysis with deployability guidelines for livestreaming. Across three standard VSR benchmarks, the proposed LiveVSR-Net framework achieves competitive performance: 23.5% WER on LRS2, 25.1% WER on LRS3, and 0.9% on GRID. By combining 3D convolutional visual encoders, Conformer-based temporal modeling, hybrid CTC-attention decoding, and domain-adapted language models, the method achieves real-time recognition suitable for deployment in production environments.