Compared to other corpus tools, such as CQPweb, Sketch Engine, and AntConc, NooJ is rarely used for collocation analysis. This might be due to the absence of an automatic collocation extraction function in NooJ. In this paper, I demonstrate that NooJ can be used for collocation analysis using a semi-automatic method. The target corpus is the Indonesian version of the Holy Quran, whose translation is approved by the Ministry of Religions of Indonesia. The target node words are all word forms derived from the lemma <lahir> with ambiguous meanings. I aim to identify the senses of node words and their collocates. The target corpus is annotated using SANTI-morf, a NooJ-based Indonesian morpheme tagger, which includes a lemmatizer. Two senses of the nodes are found: literal (SENSE_1) and metaphorical meanings (SENSE_2). Collocates are extracted semi-automatically from concordances. Thirty-one collocates are observed and lexicogrammatically categorized. The implications for disambiguation of the senses are also discussed. The results may contribute to developing SANTI-network, a NooJ-based multi-level tagger for Indonesian that combines morpheme, POS, and syntactic annotation.</lahir>

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

A Collocation Analysis of Lemma  in the Indonesian Translation of the Holy Quran

  • Prihantoro

摘要

Compared to other corpus tools, such as CQPweb, Sketch Engine, and AntConc, NooJ is rarely used for collocation analysis. This might be due to the absence of an automatic collocation extraction function in NooJ. In this paper, I demonstrate that NooJ can be used for collocation analysis using a semi-automatic method. The target corpus is the Indonesian version of the Holy Quran, whose translation is approved by the Ministry of Religions of Indonesia. The target node words are all word forms derived from the lemma  with ambiguous meanings. I aim to identify the senses of node words and their collocates. The target corpus is annotated using SANTI-morf, a NooJ-based Indonesian morpheme tagger, which includes a lemmatizer. Two senses of the nodes are found: literal (SENSE_1) and metaphorical meanings (SENSE_2). Collocates are extracted semi-automatically from concordances. Thirty-one collocates are observed and lexicogrammatically categorized. The implications for disambiguation of the senses are also discussed. The results may contribute to developing SANTI-network, a NooJ-based multi-level tagger for Indonesian that combines morpheme, POS, and syntactic annotation.