Speaker diarization, determining “who spoke what”, is a critical component of automated, speech-based educational technologies. However, in single-microphone classroom environments, there are many challenges that can make diarization difficult such as background noise, speaker overlap, and interjections. Traditional audio-only diarization methods often fail to capture the nuanced conversational dynamics of a classroom, where the linguistic patterns and discourse styles can provide important context into educator role (i.e. teacher/student). In this paper, we propose a multimodal pipeline that integrates Large Language Model (LLM) re-scoring to enhance classroom speaker diarization. We explore two improvements, the first of which involves Adaptive Centroid Enrollment (ACE), which is able to re-structure the audio embedding space using a cluster taken from the recording itself. The second, involves finetuning generative pretrained transformers to refine clustering predictions using transcriptions associated with each audio segment. Experimental results on our internal classroom dataset reveal relative improvements of 8.2% in word-level accuracy and 16.5% in word-level unweighted F1-score, highlighting the potential of LLM-based approaches for addressing speaker diarization challenges in educational settings.

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

Multimodal Classroom Diarization with GPT Re-scoring: Teacher or Student?

  • Matthew Perez,
  • Berk Coker,
  • Kemal Berk Kocabagli,
  • Jessica Vitale,
  • Alyssa Van Camp

摘要

Speaker diarization, determining “who spoke what”, is a critical component of automated, speech-based educational technologies. However, in single-microphone classroom environments, there are many challenges that can make diarization difficult such as background noise, speaker overlap, and interjections. Traditional audio-only diarization methods often fail to capture the nuanced conversational dynamics of a classroom, where the linguistic patterns and discourse styles can provide important context into educator role (i.e. teacher/student). In this paper, we propose a multimodal pipeline that integrates Large Language Model (LLM) re-scoring to enhance classroom speaker diarization. We explore two improvements, the first of which involves Adaptive Centroid Enrollment (ACE), which is able to re-structure the audio embedding space using a cluster taken from the recording itself. The second, involves finetuning generative pretrained transformers to refine clustering predictions using transcriptions associated with each audio segment. Experimental results on our internal classroom dataset reveal relative improvements of 8.2% in word-level accuracy and 16.5% in word-level unweighted F1-score, highlighting the potential of LLM-based approaches for addressing speaker diarization challenges in educational settings.