A transfer learning with data augmentation approach to emotion classification of Indonesian tweets
摘要
Emotion classification on social media provides valuable insights into public sentiment, but the performance of existing models is often limited by corpus size and linguistic variability. This research presents a transfer learning approach to analyze the benchmark EmoT corpus of 4401 emotion-labeled Indonesian Tweets, combined with a task-specific data augmentation strategy to enhance model generalization. Statistical analysis is performed using a linear mixed model of 10-fold cross validation folds, with folds modeled as a random intercept to control for within-fold variation. Results reveal that both Model and Augmentation Strategy have a significant effect on Accuracy, Macro F1, and Weighted F1 metrics. The top-performing model-strategy combination is IndoRoBERTa with augmentation via one-phase back translation, achieving a Weighted F1 score of approximately 0.859. These results highlight the effectiveness of integrating transfer learning with textual data augmentation for emotion classification in low-resource languages and suggest promising directions for future research in natural language processing.