A Collocation Analysis of Lemma in the Indonesian Translation of the Holy Quran
摘要
Compared to other corpus tools, such as CQPweb, Sketch Engine, and AntConc, NooJ is rarely used for collocation analysis. This might be due to the absence of an automatic collocation extraction function in NooJ. In this paper, I demonstrate that NooJ can be used for collocation analysis using a semi-automatic method. The target corpus is the Indonesian version of the Holy Quran, whose translation is approved by the Ministry of Religions of Indonesia. The target node words are all word forms derived from the lemma with ambiguous meanings. I aim to identify the senses of node words and their collocates. The target corpus is annotated using SANTI-morf, a NooJ-based Indonesian morpheme tagger, which includes a lemmatizer. Two senses of the nodes are found: literal (SENSE_1) and metaphorical meanings (SENSE_2). Collocates are extracted semi-automatically from concordances. Thirty-one collocates are observed and lexicogrammatically categorized. The implications for disambiguation of the senses are also discussed. The results may contribute to developing SANTI-network, a NooJ-based multi-level tagger for Indonesian that combines morpheme, POS, and syntactic annotation.