English language intelligent expression evaluation based on multimodal interactive features
摘要
In response to the issues of strong subjectivity and poor effectiveness in current English language expression evaluation, this study combines graph neural networks and time convolutional networks to extract limb and facial interaction features and their temporal sequences, and constructs an intelligent expression evaluation model based on multimodal interaction features. It was found that in the test results of the single evaluation classification model, the accuracy, F1, and Area Under the Curve (AUC) indicators of facial evaluation all reached 0.8 or above, and improved by 5.0%, 8.2%, and 4.2% respectively compared to the comparison model, showing better performance than the comparison model. The F1 and AUC indicators of limb evaluation increased by 8.8% and 3.9% respectively compared to the comparison model, which was better than the comparison model. In the test outcomes of the single item evaluation regression model, the Mean Squared Error (MSE) of facial evaluation and limb evaluation were 0.15 and 0.14, respectively, which were 7.4% and 10.8% lower than the comparison model, respectively. R2 increased by 3.6% and 17.0% compared to the comparison model, while r increased by 5.5% and 0.2%. In overall evaluation of the regression model performance, the MSE, R2, and r index values of the research model were 0.11, 0.55, and 0.77, respectively, all of which were superior to the comparison model. In classification model performance, the accuracy and AUC index values of the research model were 0.82 and 0.81, respectively, which were better than the comparison model. The outcomes denote that the proposed model can be effectively used for intelligent evaluation of English language expression, improving the language expression ability of the speaker.