The principal challenge of transcript engineering is the fact that raw transcripts of spontaneous human conversations are typically ill-suited for downstream NLP tasks. They require many processing steps to transform them into a useful format for further processing. Key procedures include segmentation of the transcript into individual sentences, diarization to identify different speakers, restoration of the true casing, and punctuation restoration for syntactic clarity. This chapter describes these tasks in detail and presents relevant tools and techniques. However, performing the above-mentioned tasks and training statistical models requires manual annotation of transcripts, a labor-intensive and expensive process. The chapter begins with the discussion of complexities associated with transcript annotation. Then, we present challenges of transcript alignment, true case restoration, and punctuation prediction. Parts of the chapter are adapted from our previous works led by Łukasz Augustyniak [2] and Piotr Żelasko [42].

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

Transcript Engineering

  • Mikołaj Morzy

摘要

The principal challenge of transcript engineering is the fact that raw transcripts of spontaneous human conversations are typically ill-suited for downstream NLP tasks. They require many processing steps to transform them into a useful format for further processing. Key procedures include segmentation of the transcript into individual sentences, diarization to identify different speakers, restoration of the true casing, and punctuation restoration for syntactic clarity. This chapter describes these tasks in detail and presents relevant tools and techniques. However, performing the above-mentioned tasks and training statistical models requires manual annotation of transcripts, a labor-intensive and expensive process. The chapter begins with the discussion of complexities associated with transcript annotation. Then, we present challenges of transcript alignment, true case restoration, and punctuation prediction. Parts of the chapter are adapted from our previous works led by Łukasz Augustyniak [2] and Piotr Żelasko [42].