Comparing Related Languages with a Fuzzy Morphism Matching Algorithm
摘要
This paper proposes a fuzzy morphism matching algorithm for discovering similarities within related languages. The fuzzy morphism matching algorithm takes as input a novel representation of the linguistic structures of the two languages that are compared. This representation is a type of Markov model that is built from an abstract representation of the basic set of words in the languages where the abstraction is based on combinations of six phoneme categories and three positions of those phonemes within the basic sets of words. The limited number of nodes in these Markov models allows efficient calculations of partial subgraph isomorphism matchings between them, and the degree of matching leads to a natural similarity measure that depends not on the number of cognate words but only on the phonetic structure of the languages, which have greater stability. This allows the detection of a strong similarity between closely related languages such as English and German as well as a weaker similarity between more distantly related languages like English and Hungarian.