Multi-label document classification for Urdu: corpus and methods
摘要
Automatically assigning various labels to language data is a core NLP task that has attracted the attention of researchers in multiple fields, for instance, medical records with diagnostic labels, products with categories, and so forth. All these tasks require an annotated multi-label corpus for supervised learning. Existing multi-label annotated corpora primarily focus on English, with a shortage in the Urdu language, which is widely spoken worldwide and whose digital text is increasing rapidly. To fill this gap, this research study presents the first large benchmark Urdu multi-label document corpus, which contains 600 documents from the field of journalism. The proposed corpus is manually annotated with 232 tags of a semantic classification scheme, with two to six tags per document to provide fine-grained categorization of the textual document. To evaluate the proposed corpus, we extracted content-based features and applied several multi-label classifiers. In addition, the results of the proposed content-based methods are compared with several deep and transfer learning-based methods. After conducting a comprehensive set of experiments, the best results are obtained using SOTA multi-label machine learning methods on word tri-gram (MicroPrecision = 0.62, MicroRecall = 0.48, MicroF1 = 0.62). To foster research in Urdu NLP, our proposed multi-label Urdu document classification corpus has been made publicly available for research and benchmarking purposes.