Understanding Sentiment in User Feedback: Lexicon-Based vs. Generative AI Approaches
摘要
Understanding feedback is vital for usability studies and user research. Incorporating sentiment analysis into qualitative data analysis methods, such as think-aloud protocols and interviews, allows researchers to gain deeper insights into users’ emotional responses that may go unnoticed. This comparative study seeks to guide researchers and practitioners in selecting the most efficient sentiment analysis tool for user experience, usability, and related fields. Two experiments were carried out to compare sentiment analyses performed by different tools. In each experiment, eleven sentiment analyses were compared–four performed with lexicon-based tools and seven with Generative AI–using two datasets of sentences manually classified by humans as positive or negative. A public dataset of 3,000 sentences from online reviews was used for the first experiment and an in-house dataset of 100 sentences from usability tests for the second one. The reliability of these tools was measured using Cohen’s Kappa coefficient to assess their agreement with human classification. Results showed that Generative AI tools outperformed lexicon-based ones. In general, most Generative AI tools show high reliability and almost perfect agreement with human classification, achieving Kappa values from 0.80 upwards in both experiments. In contrast, both experiments show that all lexicon-based tools had lower levels of agreement with human classification, achieving Kappa values between 0.21 and 0.58. Overall, the Generative AI tools proved more reliable in capturing subjectivity and contextual complexity in the language.