Developing Advanced Question-Answering Models for Legal Kazakh Texts: A Comparative Study of Modern Approaches
摘要
This study addresses the challenges of developing question-answering systems (QAS) tailored for the Kazakh legal domain. We explore modern NLP models, including GPT-3.5, Llama-2, and Llama-3, assessing their effectiveness in processing legislative texts and question-answer datasets. Leveraging structured data from two major Kazakh legal platforms, we implemented and evaluated methodologies for data scraping, processing, and model fine-tuning. The findings reveal the superior performance of the Llama-3 model in terms of accuracy, F1 score, and adaptability to the Kazakh language, despite its low-resource nature. Our results demonstrate the feasibility of advanced QAS for legal texts, paving the way for improved public access to legal information.