In recent times, there have been significant advancements in the field of LLM research, especially in the Indic LLM sphere. Sanskrit, being an integral part of India’s culture and heritage, forms the root of many other Indian languages. Current large language models often lack a deep understanding of language structure and meaning, especially for Sanskrit, leading to outputs that can be grammatically incorrect and logically flawed or biased. In this paper, we propose the Panini Sutra-based AI model that draws inspiration from Panini’s Ashtadhyayi, an ancient treatise on Sanskrit grammar consisting of nearly 4,000 sutras. To signify the need for this model, we benchmark the Indic Language Capabilities of existing Large Language Models by variating certain parameters. These include benchmarking outputs from models using different tokenizers on the same dataset and fine-tuning an existing open-source model to highlight the lack of an annotated dataset and poor linguistic capabilities. We used character-level and sub-word tokenizers like Google’s SentencePiece and OpenAI’s Tiktokenizer. The Gemma 2B model is leveraged for fine-tuning and is evaluated on the accuracy of translation between Sanskrit and English. We reckon the results obtained can broaden the dimensions and support the development of Sanskrit-based LLM.

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

Sanskrit LLM: Proving the Need for a Panini Sutra Rule-Based AI Model by Benchmarking the Capabilities of Existing LLMs

  • Shreya Srikant,
  • Vismaya Murali,
  • Varsha Badri,
  • S. S. Shylaja,
  • Mukund Rao

摘要

In recent times, there have been significant advancements in the field of LLM research, especially in the Indic LLM sphere. Sanskrit, being an integral part of India’s culture and heritage, forms the root of many other Indian languages. Current large language models often lack a deep understanding of language structure and meaning, especially for Sanskrit, leading to outputs that can be grammatically incorrect and logically flawed or biased. In this paper, we propose the Panini Sutra-based AI model that draws inspiration from Panini’s Ashtadhyayi, an ancient treatise on Sanskrit grammar consisting of nearly 4,000 sutras. To signify the need for this model, we benchmark the Indic Language Capabilities of existing Large Language Models by variating certain parameters. These include benchmarking outputs from models using different tokenizers on the same dataset and fine-tuning an existing open-source model to highlight the lack of an annotated dataset and poor linguistic capabilities. We used character-level and sub-word tokenizers like Google’s SentencePiece and OpenAI’s Tiktokenizer. The Gemma 2B model is leveraged for fine-tuning and is evaluated on the accuracy of translation between Sanskrit and English. We reckon the results obtained can broaden the dimensions and support the development of Sanskrit-based LLM.