Portable text input methods have broad application prospects in human-computer interaction systems. Especially, inertial-based gesture input attracts more attention due to its advantages of good privacy, flexibility and naturalness. However, the subtle variations in keystroke actions, personal habit differences and the sensitivity of IMUs to sensor position changes bring significant challenges for highly accurate and user-friendly gesture recognition. To solve these problems, we propose a hierarchical classification framework with spatial-temporal attention for text input using IMU gloves. Firstly, we adopt a top-down hierarchical strategy to reduce the recognition complexity by adjusting the fine-to-coarse classification paradigm. Then, a data augmentation method based on individual differences and samples variability is introduced to mitigate the effects caused by sensor position changes, action repetitions and diverse motion patterns. Afterwards, a spatial-temporal attention module is introduced to explore rich spatial-temporal information among IMUs. Our method achieves impressive recognition rates of 96.95%, 92.89% in intrasubject and intersubject scenarios, respectively. Finally, we build a text input system called GloveTyping based on self-developed gloves for online evaluation, achieving an online accuracy of 90.25%, an input speed of 7.68WPM and an error rate of 6.6%, providing more possible solution for text input.

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

GloveTyping: A Hand Gesture Recognition System for Text Input Using a Hierarchical Framework with Attention Mechanism

  • Xin Sheng,
  • Zheng Wang,
  • Haoyang Zhang,
  • Tao Zhen,
  • Pengfei Ren,
  • Liang Xie,
  • Ye Yan,
  • Erwei Yin

摘要

Portable text input methods have broad application prospects in human-computer interaction systems. Especially, inertial-based gesture input attracts more attention due to its advantages of good privacy, flexibility and naturalness. However, the subtle variations in keystroke actions, personal habit differences and the sensitivity of IMUs to sensor position changes bring significant challenges for highly accurate and user-friendly gesture recognition. To solve these problems, we propose a hierarchical classification framework with spatial-temporal attention for text input using IMU gloves. Firstly, we adopt a top-down hierarchical strategy to reduce the recognition complexity by adjusting the fine-to-coarse classification paradigm. Then, a data augmentation method based on individual differences and samples variability is introduced to mitigate the effects caused by sensor position changes, action repetitions and diverse motion patterns. Afterwards, a spatial-temporal attention module is introduced to explore rich spatial-temporal information among IMUs. Our method achieves impressive recognition rates of 96.95%, 92.89% in intrasubject and intersubject scenarios, respectively. Finally, we build a text input system called GloveTyping based on self-developed gloves for online evaluation, achieving an online accuracy of 90.25%, an input speed of 7.68WPM and an error rate of 6.6%, providing more possible solution for text input.