SaSa: Design and Implementation of a Novel Multi-category Poetry Classification Algorithm for Low-Resource Koshur Language
摘要
Poetry is an important literary genre in computational linguistics. Poetry poses a greater challenge to the use of Natural Language Processing (NLP) algorithms than any other literary genre. High degrees of accuracy are difficult to achieve in Indian languages due to their complex morphology and extensive word fusion. In this study 1719 poems in Koshur poetry are categorised into five groups namely Devotional (D), Love (L), Nature (N), Patriotic (P) and Others (O) using authors’ novel SaSa algorithm. These poems were passed through stages, tokenizing, constructing Bags-of-Words (BoW), and using TF-IDF representation, finally yielding 82,449 tokens. The various features which were extractedinclude linguistic and statistical and used to predict the input poems, while considering the count of the frequencies to remove any disambiguity. The results were validated through three annotators with an Inter Annotator Agreement (IAA) score of 88% using Cohen’s Kappa. Since no work has been done in Koshur Devnagari scripted poetry before with reference to NLP, the authors’ accuracy of 89% is regarded as good given the language's extreme lack of resourcefulness.