<p>The Tenyidie language, a.k.a, the Angami language is a low-resource language belonging to the Tibeto-Burman language family and is considered a major language of Nagaland in the north-eastern part of India. Among the many Natural Language Processing (NLP) tasks, named entity recognition (NER) is an important task in which named entities such as person, organization, location, etc, are identified and find its applications in many other applications such as classifying content for news providers, recommendation systems, sentiment analysis, etc. To the best of the authors’ knowledge, this is the first attempt at building NER for the Tenyidie language. The main aim of this research is to develop and evaluate the Named Entity Recognition (NER) annotated corpus for the Tenyidie Language. In this work, a NER annotated dataset of 10,000 sentences (211,364 tokens) for Tenyidie Language comprising of 5,208 named entities (699 persons, 2,089 organizations, and 2,420 location entities) has been created. This paper also applies the Machine Learning/Deep Learning techniques to the created NER dataset for the Tenyidie Language. For deep learning, we have explored different word embedding methods like word2vec, GloVe, fasttext, and BERT. In our experiments conducted, we achieved the best f1-score using the BERT-BASE (cased) model. The main contributions of this research are the creation of an NER annotated dataset for the Tenyidie language and the evaluation of the NER dataset using different learning techniques such as CRF, BLSTM, including the state-of-the-art BERT model.</p>

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

Tenyidie named entity recognition: corpus creation and machine/deep learning applications

  • Teisovi Angami,
  • Themrichon Tuithung

摘要

The Tenyidie language, a.k.a, the Angami language is a low-resource language belonging to the Tibeto-Burman language family and is considered a major language of Nagaland in the north-eastern part of India. Among the many Natural Language Processing (NLP) tasks, named entity recognition (NER) is an important task in which named entities such as person, organization, location, etc, are identified and find its applications in many other applications such as classifying content for news providers, recommendation systems, sentiment analysis, etc. To the best of the authors’ knowledge, this is the first attempt at building NER for the Tenyidie language. The main aim of this research is to develop and evaluate the Named Entity Recognition (NER) annotated corpus for the Tenyidie Language. In this work, a NER annotated dataset of 10,000 sentences (211,364 tokens) for Tenyidie Language comprising of 5,208 named entities (699 persons, 2,089 organizations, and 2,420 location entities) has been created. This paper also applies the Machine Learning/Deep Learning techniques to the created NER dataset for the Tenyidie Language. For deep learning, we have explored different word embedding methods like word2vec, GloVe, fasttext, and BERT. In our experiments conducted, we achieved the best f1-score using the BERT-BASE (cased) model. The main contributions of this research are the creation of an NER annotated dataset for the Tenyidie language and the evaluation of the NER dataset using different learning techniques such as CRF, BLSTM, including the state-of-the-art BERT model.