This study investigates the predictive relationship between natural hand motion dynamics and audio features in gesture-driven music systems, addressing the limitations of fixed mappings and specialized hardware. Using markerless tracking via a standard laptop camera, we capture two kinematic features of the Right Hand Index Finger: vertical position (Y-Coordinate) and speed. These features are statistically correlated with audio parameters—Fundamental Frequency ( \(f_0\) ) and RMS Loudness. Our analysis identifies two significant relationships: 1) a strong association between hand speed and loudness ( \(r = 0.535, p = 1.00 \times 10^{-3}\) ), with Granger causality confirming speed predicts loudness changes ( \(p = 1.53 \times 10^{-31}\) ); and 2) a statistically robust link between vertical position and \(f_0\) ( \(r = 0.370, p = 1.57 \times 10^{-173}\) ), supported by Granger causality ( \(p = 9.60 \times 10^{-38}\) ). These results demonstrate that natural hand motion can quantitatively predict audio parameters, providing empirical support for adaptive gesture-to-sound mapping. By leveraging ubiquitous technology, our findings enable the design of intuitive, adaptive music systems that translate natural hand gestures into sound, advancing multimodal HCI for real-time interaction.

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

Hand Motion Dynamics for Audio Prediction in Ubiquitous, Gesture-Driven Music Systems

  • Azeema Yaseen,
  • Sutirtha Chakraborty,
  • Joseph Timoney

摘要

This study investigates the predictive relationship between natural hand motion dynamics and audio features in gesture-driven music systems, addressing the limitations of fixed mappings and specialized hardware. Using markerless tracking via a standard laptop camera, we capture two kinematic features of the Right Hand Index Finger: vertical position (Y-Coordinate) and speed. These features are statistically correlated with audio parameters—Fundamental Frequency ( \(f_0\) ) and RMS Loudness. Our analysis identifies two significant relationships: 1) a strong association between hand speed and loudness ( \(r = 0.535, p = 1.00 \times 10^{-3}\) ), with Granger causality confirming speed predicts loudness changes ( \(p = 1.53 \times 10^{-31}\) ); and 2) a statistically robust link between vertical position and \(f_0\) ( \(r = 0.370, p = 1.57 \times 10^{-173}\) ), supported by Granger causality ( \(p = 9.60 \times 10^{-38}\) ). These results demonstrate that natural hand motion can quantitatively predict audio parameters, providing empirical support for adaptive gesture-to-sound mapping. By leveraging ubiquitous technology, our findings enable the design of intuitive, adaptive music systems that translate natural hand gestures into sound, advancing multimodal HCI for real-time interaction.