<p>Morphological analysis is essential for natural language processing (NLP), particularly for low-resource languages like Dogri, which lacks standardized linguistic resources. The primary contribution lies in developing the first Bi-LSTM-based morphological analyzer for Dogri that predicts root words and grammatical features with high accuracy. It evaluates two label representation techniques: monolithic and individual feature representation, to determine their impact on accuracy. A dataset of 290,361 Dogri words was manually annotated with morpheme boundaries and grammatical features such as gender, number, case, and tense. The Bi-LSTM model was trained using character embeddings, and its performance was compared against Recurrent Neural Networks (RNN) and traditional paradigm-based approaches. The results show significant accuracy improvements across all Parts of Speech (POS) categories. The individual feature representation technique outperformed the monolithic representation, achieving high accuracy for nouns (98.03%), adjectives (96.75%), adverbs (95.41%), pronoun (82.39%) and verbs (80.01%). The Bi-LSTM model demonstrated superior performance over RNN and paradigm-based methods. However, challenges remain in predicting complex inflectional patterns, particularly in pronouns and verbs. This research not only addresses a critical gap in Dogri NLP resources but also establishes a foundational framework for scalable morphological analysis applicable to other morphologically rich, low-resource languages. Future work will focus on incorporating sentence-level dependencies and hybrid rule-based approaches to improve performance.</p>

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

Enhancing NLP for Low-Resource Language by Developing Deep Learning-Powered Morphological Analysis of Dogri: An End-to-End Pipeline from Corpus Construction and Linguistic Annotation to Model Training and Deployment

  • Parul Gupta,
  • Shubhnandan Singh Jamwal

摘要

Morphological analysis is essential for natural language processing (NLP), particularly for low-resource languages like Dogri, which lacks standardized linguistic resources. The primary contribution lies in developing the first Bi-LSTM-based morphological analyzer for Dogri that predicts root words and grammatical features with high accuracy. It evaluates two label representation techniques: monolithic and individual feature representation, to determine their impact on accuracy. A dataset of 290,361 Dogri words was manually annotated with morpheme boundaries and grammatical features such as gender, number, case, and tense. The Bi-LSTM model was trained using character embeddings, and its performance was compared against Recurrent Neural Networks (RNN) and traditional paradigm-based approaches. The results show significant accuracy improvements across all Parts of Speech (POS) categories. The individual feature representation technique outperformed the monolithic representation, achieving high accuracy for nouns (98.03%), adjectives (96.75%), adverbs (95.41%), pronoun (82.39%) and verbs (80.01%). The Bi-LSTM model demonstrated superior performance over RNN and paradigm-based methods. However, challenges remain in predicting complex inflectional patterns, particularly in pronouns and verbs. This research not only addresses a critical gap in Dogri NLP resources but also establishes a foundational framework for scalable morphological analysis applicable to other morphologically rich, low-resource languages. Future work will focus on incorporating sentence-level dependencies and hybrid rule-based approaches to improve performance.