In this paper, we have shown the development of a Part of Speech (POS) tagger for Hadoti - a prominent language spoken in Rajasthan, India - despite its limited resources. For this, we manually tagged a corpus of 50,000 POS-tagged sentences and trained it using a Hidden Markov Model (HMM). Since no prior work had been reported in this area, we couldn't compare our results to any other system. This paper documents the efforts made to create an HMM POS tagger for Hadoti, to stimulate further research in this field. This work is expected to serve as a foundation for preserving the language and as a resource for aspiring researchers who wish to explore this area of Hadoti Language Processing. The system was evaluated for accuracy and produced 99.87% accurate results on seen data and 98.78% on unseen data. The system was able to produce an accuracy of 99.33% on the entire test corpus.

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

Development of HMM Based Parts of Speech Tagger for Hadoti

  • Anushka Nagar,
  • Nisheeth Joshi,
  • Pragya Katyayan,
  • Palak Arora

摘要

In this paper, we have shown the development of a Part of Speech (POS) tagger for Hadoti - a prominent language spoken in Rajasthan, India - despite its limited resources. For this, we manually tagged a corpus of 50,000 POS-tagged sentences and trained it using a Hidden Markov Model (HMM). Since no prior work had been reported in this area, we couldn't compare our results to any other system. This paper documents the efforts made to create an HMM POS tagger for Hadoti, to stimulate further research in this field. This work is expected to serve as a foundation for preserving the language and as a resource for aspiring researchers who wish to explore this area of Hadoti Language Processing. The system was evaluated for accuracy and produced 99.87% accurate results on seen data and 98.78% on unseen data. The system was able to produce an accuracy of 99.33% on the entire test corpus.