<p>Lipreading is a challenging visual understanding task that operates independently of audio signals. Its core difficulty lies in the subtle contraction and relaxation of lip muscles and the rapid transitions in lip shapes, which make fine-grained lip movements hard to capture. Traditional methods primarily focus on extracting static textures and dynamic lip shape changes in the temporal domain, often neglecting the frequency characteristics of lip movements in the frequency domain–such as the periodicity of mouth opening and closing and the frequency distribution of muscle movements. This leads to suboptimal modeling of subtle lip motions. Inspired by the principle from spectral analysis that low-frequency components are more suitable for modeling subtle movements, we propose a novel lipreading framework based on temporal-frequency dual-domain feature fusion and lip motion magnification, termed D2M2Lip. Specifically, we first design a temporal feature extraction module (TFEM) to capture the dynamic sequential features from RGB video frames. Then, we introduce a frequency feature extraction module (FFEM) comprising frequency separation submodule (FSM) and lip motion magnification submodule (LMM): the former separates low-frequency components representing motion trends from high-frequency components capturing texture and edges, while the latter magnifies low-frequency components associated with subtle lip movements. Finally, we propose a dual-domain feature fusion module that employs bidirectional cross-attention to integrate temporal and frequency features, enhancing collaborative representation and compensating for the limitations of single-domain modeling, thereby improving the robustness of subtle lip motion recognition. Extensive experiments on the CMLR and GRID datasets demonstrate that the proposed method significantly reduces the lipreading error rate, validating its effectiveness.</p>

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

D2M2Lip: dual-domain feature fusion and motion magnification for lip reading

  • Baochao Zhu,
  • Shujie Li,
  • Jinrui Zhang,
  • Feng Xue

摘要

Lipreading is a challenging visual understanding task that operates independently of audio signals. Its core difficulty lies in the subtle contraction and relaxation of lip muscles and the rapid transitions in lip shapes, which make fine-grained lip movements hard to capture. Traditional methods primarily focus on extracting static textures and dynamic lip shape changes in the temporal domain, often neglecting the frequency characteristics of lip movements in the frequency domain–such as the periodicity of mouth opening and closing and the frequency distribution of muscle movements. This leads to suboptimal modeling of subtle lip motions. Inspired by the principle from spectral analysis that low-frequency components are more suitable for modeling subtle movements, we propose a novel lipreading framework based on temporal-frequency dual-domain feature fusion and lip motion magnification, termed D2M2Lip. Specifically, we first design a temporal feature extraction module (TFEM) to capture the dynamic sequential features from RGB video frames. Then, we introduce a frequency feature extraction module (FFEM) comprising frequency separation submodule (FSM) and lip motion magnification submodule (LMM): the former separates low-frequency components representing motion trends from high-frequency components capturing texture and edges, while the latter magnifies low-frequency components associated with subtle lip movements. Finally, we propose a dual-domain feature fusion module that employs bidirectional cross-attention to integrate temporal and frequency features, enhancing collaborative representation and compensating for the limitations of single-domain modeling, thereby improving the robustness of subtle lip motion recognition. Extensive experiments on the CMLR and GRID datasets demonstrate that the proposed method significantly reduces the lipreading error rate, validating its effectiveness.