The compilation of corpora is a fundamental procedure in linguistic research. This chapter elucidates the systematic approaches undertaken in the compilation and annotation of the IRANDOC specialized corpora. It also delves into the intricacies and challenges encountered during the preprocessing stages, particularly normalization and tokenization, as well as the tagging phase. The objective of this chapter is to familiarize readers with the procedural steps and potential challenges inherent in the development of a Persian corpus. The chapter culminates with a discussion on the utility of a dedicated corpus website, which facilitates the contextual study of linguistic elements (target words) within the corpus.

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

Creating and Annotating Three Specialized Corpora in Persian: Steps and Challenges

  • Elham Alayiaboozar

摘要

The compilation of corpora is a fundamental procedure in linguistic research. This chapter elucidates the systematic approaches undertaken in the compilation and annotation of the IRANDOC specialized corpora. It also delves into the intricacies and challenges encountered during the preprocessing stages, particularly normalization and tokenization, as well as the tagging phase. The objective of this chapter is to familiarize readers with the procedural steps and potential challenges inherent in the development of a Persian corpus. The chapter culminates with a discussion on the utility of a dedicated corpus website, which facilitates the contextual study of linguistic elements (target words) within the corpus.