<p>HGR has great applications in human-computer interaction, human-robot interaction, supporting the deaf and mute, etc. To exploit the performance of DL for the problem of HGR in building a device control system, it is necessary to choose a suitable model. In this paper, we publish a dataset (TQU-HG dataset) of large RGB images of hand gestures collected with low resolution (640<InlineEquation ID="IEq1"> <InlineMediaObject> <ImageObject Color="BlackWhite" FileRef="11042_2025_20743_Article_IEq1.gif" Format="GIF" Height="13" Rendition="HTML" Resolution="72" Type="Linedraw" Width="19" /> </InlineMediaObject> <EquationSource Format="TEX">\(\times \)</EquationSource> <EquationSource Format="MATHML"><math> <mo>×</mo> </math></EquationSource> </InlineEquation>480) pixels, low light conditions, and fast collection speed (16 fps). TQU-HG dataset includes 60,000 images collected from 20 people (10 male, 10 female) with 15 gestures of both left and right hands. From there, we propose a taxonomy of comparative research on HGR model that can meet the requirements of practical applications based on the selection and comparison based on three branches on the RGB images. (1) some traditional machine learning methods combined with Mediapipe_TML to recognize hand gestures, (2) fine-tuning the HGR model based on the CNNs (YOLOv5, YOLOv6, YOLOv7, YOLOv8, YOLO-Nas, YOLOv9, YOLOv10, YOLOv11, SSD_VGG16, ResNet18, ResNet50, ResNet152, ResNext50, MobileNetV3-small, MobileNetV3-large, SSD-VGG16), (3) fine-tuning the HGR model based on the transformer-based methods (TBN and MVTN). To select the best model in comparative studies, we fine-tune it and evaluate it on TQU-HG, and HaGRID datasets. In particular, we performed cross-evaluation between TQU-HG and HaGRID datasets to select the best model with the ResNext50 method (<InlineEquation ID="IEq2"> <InlineMediaObject> <ImageObject Color="BlackWhite" FileRef="11042_2025_20743_Article_IEq2.gif" Format="GIF" Height="18" Rendition="HTML" Resolution="72" Type="Linedraw" Width="321" /> </InlineMediaObject> <EquationSource Format="TEX">\(P=99.37\%, R=99.37\%, F1Score=99.36\%\)</EquationSource> <EquationSource Format="MATHML"><math> <mrow> <mi>P</mi> <mo>=</mo> <mn>99.37</mn> <mo>%</mo> <mo>,</mo> <mi>R</mi> <mo>=</mo> <mn>99.37</mn> <mo>%</mo> <mo>,</mo> <mi>F</mi> <mn>1</mn> <mi>S</mi> <mi>c</mi> <mi>o</mi> <mi>r</mi> <mi>e</mi> <mo>=</mo> <mn>99.36</mn> <mo>%</mo> </mrow> </math></EquationSource> </InlineEquation>) on TQU-HG dataset and with the YOLO-Nas method (<InlineEquation ID="IEq3"> <InlineMediaObject> <ImageObject Color="BlackWhite" FileRef="11042_2025_20743_Article_IEq3.gif" Format="GIF" Height="18" Rendition="HTML" Resolution="72" Type="Linedraw" Width="321" /> </InlineMediaObject> <EquationSource Format="TEX">\(P=99.39\%, R=99.05\%, F1Score=99.37\%\)</EquationSource> <EquationSource Format="MATHML"><math> <mrow> <mi>P</mi> <mo>=</mo> <mn>99.39</mn> <mo>%</mo> <mo>,</mo> <mi>R</mi> <mo>=</mo> <mn>99.05</mn> <mo>%</mo> <mo>,</mo> <mi>F</mi> <mn>1</mn> <mi>S</mi> <mi>c</mi> <mi>o</mi> <mi>r</mi> <mi>e</mi> <mo>=</mo> <mn>99.37</mn> <mo>%</mo> </mrow> </math></EquationSource> </InlineEquation>) on HaGRID dataset for HGR. Cross-evaluation of YOLOv9, YOLOv10, YOLOv11, and transformer-based models on TQU-HG and HaGRID datasets with the same eight labels results in <i>P</i> greater than 98%. The quantitative results of the training and testing process are presented in detail and available.</p>

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

Selected hand gesture recognition model based on cross-evaluation of deep learning from large RGB image datasets

  • Van-Hung Le

摘要

HGR has great applications in human-computer interaction, human-robot interaction, supporting the deaf and mute, etc. To exploit the performance of DL for the problem of HGR in building a device control system, it is necessary to choose a suitable model. In this paper, we publish a dataset (TQU-HG dataset) of large RGB images of hand gestures collected with low resolution (640 \(\times \) × 480) pixels, low light conditions, and fast collection speed (16 fps). TQU-HG dataset includes 60,000 images collected from 20 people (10 male, 10 female) with 15 gestures of both left and right hands. From there, we propose a taxonomy of comparative research on HGR model that can meet the requirements of practical applications based on the selection and comparison based on three branches on the RGB images. (1) some traditional machine learning methods combined with Mediapipe_TML to recognize hand gestures, (2) fine-tuning the HGR model based on the CNNs (YOLOv5, YOLOv6, YOLOv7, YOLOv8, YOLO-Nas, YOLOv9, YOLOv10, YOLOv11, SSD_VGG16, ResNet18, ResNet50, ResNet152, ResNext50, MobileNetV3-small, MobileNetV3-large, SSD-VGG16), (3) fine-tuning the HGR model based on the transformer-based methods (TBN and MVTN). To select the best model in comparative studies, we fine-tune it and evaluate it on TQU-HG, and HaGRID datasets. In particular, we performed cross-evaluation between TQU-HG and HaGRID datasets to select the best model with the ResNext50 method ( \(P=99.37\%, R=99.37\%, F1Score=99.36\%\) P = 99.37 % , R = 99.37 % , F 1 S c o r e = 99.36 % ) on TQU-HG dataset and with the YOLO-Nas method ( \(P=99.39\%, R=99.05\%, F1Score=99.37\%\) P = 99.39 % , R = 99.05 % , F 1 S c o r e = 99.37 % ) on HaGRID dataset for HGR. Cross-evaluation of YOLOv9, YOLOv10, YOLOv11, and transformer-based models on TQU-HG and HaGRID datasets with the same eight labels results in P greater than 98%. The quantitative results of the training and testing process are presented in detail and available.