Development of Telugu Treebank Using Panini Karaka Relations for Dependency Parsing
摘要
A Treebank is an annotated corpus and it is a very useful linguistic resource for building dependency parsers, word sense disambiguation systems, and question answering systems. Most of the treebanks are now available in the English language, and a very limited number of Treebanks are available in Indian languages. English and other languages follows western grammar formalism, but these grammar rules are not applicable to Indian languages. Telugu is one of the south Indian languages and the most predominant spoken language in India. This paper describes the steps for developing Telugu Treebank using Panini karaka and non-karaka relations. The developed Telugu treebank is applied to various dependency parsers. This paper also discusses the results achieved with the various dependency parsers using Telugu Treebank.