Human and Machine Keyphrase Perception in Russian Text and Speech
摘要
The article examines the perception and extraction of keyphrases in both written and spoken text. Experiments were performed on the dataset including transcripts and audio recordings of lectures by Russian-speaking participants of the project “Postnauka”. The results show that automated methods for keyphrase extraction have limited accuracy, with statistical algorithms performing the worst and generative AI models, such as ChatGPT, showing a closer resemblance to human perception. Additionally, while there is some overlap between keyphrases extracted from written and oral texts, spoken text presents greater variability. Experiments using synthesized speech indicate that listeners rely heavily on content, rather than acoustic cues, when understanding spoken text. Acoustic analysis reveals that keyphrases are distinguished by longer duration, wider pitch range, and higher energy, aligning with previous findings in other languages.