The Sanskrit Sembank
摘要
This paper introduces the Sanskrit Sembank (ssb), a comprehensive lexical semantic resource integrated within the Digital Corpus of Sanskrit. The ssb combines lexicographic data from Sanskrit dictionaries with synsets derived from the Princeton WordNet, and offers lexical semantic annotations of more than 600,000 words in context across both Vedic and Classical Sanskrit texts. We discuss the methodological challenges in adapting WordNet’s conceptual framework to the vocabulary of Sanskrit, particularly in religious and scientific domains. We further present evaluation results from word sense disambiguation experiments, achieving F1 scores of up to 86.7% without using an LLM. The paper also examines the complementary relationship between the ssb and the Sanskrit WordNet project, highlighting their distinct and complementary approaches to lexical semantic annotation.