Emotion in Context: Human and AI Perspectives on the Affective Landscape of Severance
摘要
Emotion detection has long been a focus of interdisciplinary research spanning psychology, linguistics, and discourse analysis. Traditional computational approaches to emotion detection have included both lexicon-based and machine learning-based methods. While lexicon-based approaches offer interpretability and domain independence, they often struggle to capture contextual nuance in complex narratives. Machine learning methods, though effective within their training domains, require substantial annotated data and face challenges with domain adaptation. Recent advances in natural language processing (NLP), particularly the development of large language models (LLMs), offer new possibilities for automating emotion detection with greater contextual awareness. This study investigates how emotions are represented and perceived in the television series Severance (Apple TV, 2022) through a multi-layered annotation framework based on Parrott’s hierarchical taxonomy of emotions. All nine episodes of season 1 were annotated scene by scene, producing four datasets: human text-only, human multimodal (video + audio + text), and two LLMs. Comparative analyses assessed inter-annotator agreement, modality effects, and temporal and character-based emotion patterns. Results show that multimodal annotations yield a far richer and more nuanced emotional landscape than text alone, highlighting the decisive role of prosody, gesture, and facial expression in conveying affect. LLMs captured general emotional polarity but failed to reproduce fine-grained distinctions. Character profiles further reveal individualized affective strategies of repression, resistance, and awakening. The findings underscore the inherently multimodal nature of emotional meaning in audiovisual narrative.