<p>Using a corpus of 264,851 Chinese-language pop song lyrics spanning 1967–2023, this study introduces a two-step sentiment-emotion analysis pipeline that couples the generative synthesis of ChatGPT with the rule-based transparency of a 38,237-entry Chinese emotion lexicon. ChatGPT first condenses every lyric into its three most salient affective cues. These cues are then quantified, via the lexicon, across a four-level emotion hierarchy comprising (1) a global sentiment, (2) positive and negative intensities, (3) eight primary emotions, and (4) 23 sub-emotions. Preliminary evaluation indicates that this hybrid method attains 87% accuracy on predicting binary sentiment (positive/negative) of lyrics, outperforming a lexicon-only baseline by 23 points while retaining interpretability. Temporal analysis uncovers a pronounced 35-year affective cycle that crests in the late-1970s and then around 2010, and a trough in the late-1980s. Opposite shifts in <i>joy</i> and <i>sadness</i>, and secondarily in <i>liking</i> and <i>disgust</i>, drive this cycle, while emotional richness and conflictedness peak when negativity is high, revealing periods of densely layered ambivalence. K-means clustering groups songs into four archetypal palettes—“resentful heartbreak,” “happy romance,” “bittersweet love,” and “sad romance”—underscoring the centrality of affection as Mandopop’s emotional glue. Network analysis further identifies emotions <i>liking</i> and <i>sadness</i>, and sub-emotions “fondness,” “sorrow,” and “annoyance,” as the chief bridges knitting positive and negative affect, enabling complex emotional tapestries across the corpus. By marrying large-language-model insight with lexicon consistency, the framework delivers a scalable, fine-grained, and partially replicable method for charting textual emotion dynamics, offering new avenues for comparative digital humanities and affective cultural analytics.</p>

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

Mapping the emotion-scape in Chinese-language pop song lyrics, 1967–2023: combining LLM with lexicon-based sentiment analysis

  • Xiaolu Wang,
  • Evan Wong

摘要

Using a corpus of 264,851 Chinese-language pop song lyrics spanning 1967–2023, this study introduces a two-step sentiment-emotion analysis pipeline that couples the generative synthesis of ChatGPT with the rule-based transparency of a 38,237-entry Chinese emotion lexicon. ChatGPT first condenses every lyric into its three most salient affective cues. These cues are then quantified, via the lexicon, across a four-level emotion hierarchy comprising (1) a global sentiment, (2) positive and negative intensities, (3) eight primary emotions, and (4) 23 sub-emotions. Preliminary evaluation indicates that this hybrid method attains 87% accuracy on predicting binary sentiment (positive/negative) of lyrics, outperforming a lexicon-only baseline by 23 points while retaining interpretability. Temporal analysis uncovers a pronounced 35-year affective cycle that crests in the late-1970s and then around 2010, and a trough in the late-1980s. Opposite shifts in joy and sadness, and secondarily in liking and disgust, drive this cycle, while emotional richness and conflictedness peak when negativity is high, revealing periods of densely layered ambivalence. K-means clustering groups songs into four archetypal palettes—“resentful heartbreak,” “happy romance,” “bittersweet love,” and “sad romance”—underscoring the centrality of affection as Mandopop’s emotional glue. Network analysis further identifies emotions liking and sadness, and sub-emotions “fondness,” “sorrow,” and “annoyance,” as the chief bridges knitting positive and negative affect, enabling complex emotional tapestries across the corpus. By marrying large-language-model insight with lexicon consistency, the framework delivers a scalable, fine-grained, and partially replicable method for charting textual emotion dynamics, offering new avenues for comparative digital humanities and affective cultural analytics.