E-commerce video & content understanding via vision: a comprehensive study on lip reading and visual speech recognition for enhanced shopping experience
摘要
Livestream e-commerce now exceeds USD 500 billion in annual online retail sales. Nevertheless, a large proportion of viewers watch livestreams without sound, reducing engagement and lowering purchase conversion rates. In this work, we explore vision-based content understanding through lip reading and visual speech recognition (VSR) for automatic subtitle generation to enhance livestream clarity. We present LiveVSR-Net, a novel architecture combining 3D convolutional neural networks with transformer-based temporal modeling and hybrid CTC-attention decoding. Domain-adapted enhancements include e-commerce-specific vocabulary expansion, fine-tuned language model integration, and multi-speaker handling techniques. We conduct comprehensive ablation studies to quantify the contribution of each architectural component, evaluate on standard benchmarks (LRS2, LRS3, GRID), and provide real-world robustness analysis with deployability guidelines for livestreaming. Across three standard VSR benchmarks, the proposed LiveVSR-Net framework achieves competitive performance: 23.5% WER on LRS2, 25.1% WER on LRS3, and 0.9% on GRID. By combining 3D convolutional visual encoders, Conformer-based temporal modeling, hybrid CTC-attention decoding, and domain-adapted language models, the method achieves real-time recognition suitable for deployment in production environments.