Santhali is one of the popular local languages and is mainly spoken by tribal people of States like Jharkhand, Odisha, and West Bengal, as well as a few other parts of the country. Tokenization is one of the major challenges in Natural Language Processing (NLP). Tokenization reduces the length of sentences and paragraphs. There are various methods in use to tokenize different languages depending on the writing and structure of the language. Tokenization is important in many tasks like machine translation, text summarization, chatbot creation, and spam detection. In Natural Language Processing (NLP) most of the Indian spoken languages have high resource availability and have improved technologically in various fields. Still, a low-resource language like SANTHALI does not make much of its existence in technology. The present work is based on the tokenization of Santhali. All available tokenizers to tokenize in the Santhali language have been explored. Ol-Chiki font has been used for the present study. A detailed comparison of tokenization has been represented in the manuscript in the present study. It has been observed that for Santhali a writing system is alphabetic and most of the characters are made of real-world symbols. spaCy model performs better than other available tokenizer models based on the mechanism of the Santhali language.

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

Tokenization's Potential for Santhali Language Processing

  • Anand Kumar Ohm,
  • Aman Kumar Panday,
  • Koushlendra Kumar Singh

摘要

Santhali is one of the popular local languages and is mainly spoken by tribal people of States like Jharkhand, Odisha, and West Bengal, as well as a few other parts of the country. Tokenization is one of the major challenges in Natural Language Processing (NLP). Tokenization reduces the length of sentences and paragraphs. There are various methods in use to tokenize different languages depending on the writing and structure of the language. Tokenization is important in many tasks like machine translation, text summarization, chatbot creation, and spam detection. In Natural Language Processing (NLP) most of the Indian spoken languages have high resource availability and have improved technologically in various fields. Still, a low-resource language like SANTHALI does not make much of its existence in technology. The present work is based on the tokenization of Santhali. All available tokenizers to tokenize in the Santhali language have been explored. Ol-Chiki font has been used for the present study. A detailed comparison of tokenization has been represented in the manuscript in the present study. It has been observed that for Santhali a writing system is alphabetic and most of the characters are made of real-world symbols. spaCy model performs better than other available tokenizer models based on the mechanism of the Santhali language.