Transcript Engineering
摘要
The principal challenge of transcript engineering is the fact that raw transcripts of spontaneous human conversations are typically ill-suited for downstream NLP tasks. They require many processing steps to transform them into a useful format for further processing. Key procedures include segmentation of the transcript into individual sentences, diarization to identify different speakers, restoration of the true casing, and punctuation restoration for syntactic clarity. This chapter describes these tasks in detail and presents relevant tools and techniques. However, performing the above-mentioned tasks and training statistical models requires manual annotation of transcripts, a labor-intensive and expensive process. The chapter begins with the discussion of complexities associated with transcript annotation. Then, we present challenges of transcript alignment, true case restoration, and punctuation prediction. Parts of the chapter are adapted from our previous works led by Łukasz Augustyniak [2] and Piotr Żelasko [42].