Language is a primary means of communication. It is a medium through which we can interact with society. Recognizing it, each language has its own set of grammatical rules. This study focused on the development of a rule-based chunker for a resource-poor language Sindhi using the Devanagari script. We have chosen a rule-based approach as language itself is a sequence of rules. This approach is fairly useful due to its ability to capture nuances of a language. The language rules were created and validated with the help of language experts. To develop the chunker, 50,000 sentences were used for the construction of rules. These sentences belonged to various domains like travel and tourism, health and administration. For this study, there was a requirement for a POS-tagged dataset. The data was annotated using Part of Speech tagger based on the Hidden Markov Model (HMM). 1,000 sentences were used to evaluate the system. The developed chunker showed an accuracy of 97.8%.

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

Development of Rule-Based Chunker for Sindhi

  • Palak Arora,
  • Bharti Nathani,
  • Nisheeth Joshi,
  • Pragya Katyayan

摘要

Language is a primary means of communication. It is a medium through which we can interact with society. Recognizing it, each language has its own set of grammatical rules. This study focused on the development of a rule-based chunker for a resource-poor language Sindhi using the Devanagari script. We have chosen a rule-based approach as language itself is a sequence of rules. This approach is fairly useful due to its ability to capture nuances of a language. The language rules were created and validated with the help of language experts. To develop the chunker, 50,000 sentences were used for the construction of rules. These sentences belonged to various domains like travel and tourism, health and administration. For this study, there was a requirement for a POS-tagged dataset. The data was annotated using Part of Speech tagger based on the Hidden Markov Model (HMM). 1,000 sentences were used to evaluate the system. The developed chunker showed an accuracy of 97.8%.