Sentiment Analysis of Twitter comes mainly under the domain of Natural Language Processing, and has some crucial applications in brand monitoring, market research, customer support, and others. This research makes an extensive analysis of the performance of various machine learning models on a massive Twitter sentiment analysis repository of 1.6 million tweets. Six models—Logistic Regression, Naive Bayes, Gradient Boosting, Support Vector Machine, Random Forest, and CNN—are evaluated based on performance metrics such as accuracy, F1-score, precision, recall, and ROC–AUC. Targeting the restriction of large dataset, sampling techniques are applied to obtain an optimal pilot study sample for each model in order to maintain the best level of accuracy. Results show that the SVM records 74% accuracy and ROC-AUC of 0.81 but takes longer in terms of execution time. Nevertheless, this research introduces a novel hybrid set of classifying models that integrates both traditional and advanced deep learning techniques in order to improve the level of accuracy, level of sensitivity and specificity. This innovative hybrid approach yields an accuracy of 80%, sensitivity 81%, and ROC-AUC of 0.88, outperforming existing models. This study contributes to the advancement of Twitter sentiment analysis and recommends a cutting-edge model for improved performance. The proposed hybrid model leverages the advantages of both architectures, capturing contextual relationships while requiring fewer computational resources than deep learning models, making it a promising approach for various applications, including customer service, market research, and political analysis.

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

Comparative Analysis of Machine Learning Models for Twitter Sentiment Analysis: A Recommendation of an Inventive Approach

  • Malay Bandyapadhyay

摘要

Sentiment Analysis of Twitter comes mainly under the domain of Natural Language Processing, and has some crucial applications in brand monitoring, market research, customer support, and others. This research makes an extensive analysis of the performance of various machine learning models on a massive Twitter sentiment analysis repository of 1.6 million tweets. Six models—Logistic Regression, Naive Bayes, Gradient Boosting, Support Vector Machine, Random Forest, and CNN—are evaluated based on performance metrics such as accuracy, F1-score, precision, recall, and ROC–AUC. Targeting the restriction of large dataset, sampling techniques are applied to obtain an optimal pilot study sample for each model in order to maintain the best level of accuracy. Results show that the SVM records 74% accuracy and ROC-AUC of 0.81 but takes longer in terms of execution time. Nevertheless, this research introduces a novel hybrid set of classifying models that integrates both traditional and advanced deep learning techniques in order to improve the level of accuracy, level of sensitivity and specificity. This innovative hybrid approach yields an accuracy of 80%, sensitivity 81%, and ROC-AUC of 0.88, outperforming existing models. This study contributes to the advancement of Twitter sentiment analysis and recommends a cutting-edge model for improved performance. The proposed hybrid model leverages the advantages of both architectures, capturing contextual relationships while requiring fewer computational resources than deep learning models, making it a promising approach for various applications, including customer service, market research, and political analysis.