An Effective Deep Learning Model for Air-Scripted Alphabet Recognition System
摘要
Setting out on a quest to redefine human-computer interaction, our project prioritized the development of an innovative CNN model customized for Air-scripted alphabet recognition. Grounded in a tailor-made dataset, our methodology began with the creation of a bespoke CNN that achieved 97.23% accuracy. Subsequent refinement introduced feature extraction via VGG19, feeding into various machine learning models and a Fully Connected Neural Network (FCNN), yielding 96% accuracy. Notably, the integration of global second-order pooling (GSOP) in the FCNN framework showcased enhanced performance. Advantages of GSOP over bespoke CNN include improved computational efficiency and robust feature representation. Comprehensive evaluations, including classification reports, confusion matrix and accuracy-loss curves, underscore the superiority of our proposed GSOP-enhanced model. This research heralds a transformative approach to gesture-based alphabet recognition, promising more efficient and accurate human-computer interaction paradigms.