Automatic Grammatical Disambiguation in Belarusian and Russian Legal Domain
摘要
The article presents a parallel Russian-Belarusian corpus compiled from the Republic of Belarus law codes. We have analyzed a sub-corpus of almost sixty thousand words in detail and studied grammatical ambiguity occurring in it. We have developed rule-based and statistical methods to clear up ambiguity in Belarusian, including nineteen local grammars which solve homonymy for the twenty most frequent words and for several grammatical categories. These local grammars are used to select the correct tag among variants (nominative or accusative case, plural or singular, numeral or verb, etc.) and determine the proper case form after prepositions governing several cases.