NER for Albanian Language: A Manually Annotated Corpus and Machine Learning Models
摘要
Recent advancements in artificial intelligence (AI) have significantly enhanced tasks like named entity recognition (NER), enabling the identification of people, organizations, products, events, places, and dates. This paper introduces an Albanian NER corpus with 1,003,836 tokens (56,595 sentences), including 89,850 labelled tokens, annotated with 10 NER tags. The corpus, sourced from well-known Albanian news platforms, are used to train and evaluate 10 models using algorithms such as Naïve Bayes, Logistic Regression, SVM, Random Forest, Gradient Boosting, Extreme Gradient Boosting, and Multi-Layer Perceptron variants. Among these, Extra Trees and Random Forest achieved the best results, with approximately 96% accuracy and a 95% F1 score. As the largest NER corpus in Albanian, this resource advances linguistic research and AI applications, enhancing NER tasks and advance natural language processing (NLP) developments for the Albanian language.