Monitoring collaborative learning in educational settings is challenging, because teachers cannot track all student groups in detail. Recently, numerous works have shown that collaborative learning can be automatically tracked with machine learning. In this work, we explore how Google and Whispers automatic transcription systems influence Collaborative Problem Solving (CPS) classification performance using a public dataset. We also look at how different styles of transcribing overlapping speech influence CPS classification performance. The transcription styles in our analysis are active speaker (transcriptions of a single active speaker) and temporal (transcriptions of all participants in speaking order). As a baseline, manually segmented speech utterances were transcribed for each transcription style. We first evaluated Automatic Speech Recognition (ASR) performance as it pertains to each type. Then, we evaluate how each type influences CPS classification. We find that active speaker transcription equates with higher CPS classification performance, but that Google ASR is closer to temporal transcriptions. Our results provide insights into how to capitalize on ASRs in collaborative contexts.

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

Investigating Automated Transcriptions for Multimodal CPS Detection in Groupwork

  • Benjamin Ibarra,
  • Brett Wisniewski,
  • Corbyn Terpstra,
  • Videep Venkatesha,
  • Mariah Bradford,
  • Nathaniel Blanchard

摘要

Monitoring collaborative learning in educational settings is challenging, because teachers cannot track all student groups in detail. Recently, numerous works have shown that collaborative learning can be automatically tracked with machine learning. In this work, we explore how Google and Whispers automatic transcription systems influence Collaborative Problem Solving (CPS) classification performance using a public dataset. We also look at how different styles of transcribing overlapping speech influence CPS classification performance. The transcription styles in our analysis are active speaker (transcriptions of a single active speaker) and temporal (transcriptions of all participants in speaking order). As a baseline, manually segmented speech utterances were transcribed for each transcription style. We first evaluated Automatic Speech Recognition (ASR) performance as it pertains to each type. Then, we evaluate how each type influences CPS classification. We find that active speaker transcription equates with higher CPS classification performance, but that Google ASR is closer to temporal transcriptions. Our results provide insights into how to capitalize on ASRs in collaborative contexts.