The Google Books Ngram diachronic corpus is an important tool to study changes in the language and society. The article draws attention to the problem of diachronic balance instability of various text genres in the corpus. The instability is shown not to correspond the publishing policy of the country. The method for detecting the imbalance is proposed. The Russian subcorpus is analyzed in detail to shows the presence of serious balance instabilities in the 21st century. We analyze the frequency dynamics of small groups of words, as well as make extended statistical research of the corpus. We calculate such parameters of word frequency dynamics as autocorrelations, linear trends, mathematical expectations and variances. We also apply the methods of singular spectral analysis and moving average to find unnatural peaks of the average “noise”. Our conclusion is unreasonably large amount of fiction in the Russian subcorpus in some specific years. Similar effects are also found in the English subcorpus. These facts seem to be significant because the most often studied objects of psychological and sociological research are the high-frequency common words widespread in fiction. An example is given to demonstrate that ignoring this factor can result in erroneous conclusions.

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

How to Detect Imbalances in the Google Books Ngram Corpus?

  • Valery Solovyev,
  • Anna Ivleva

摘要

The Google Books Ngram diachronic corpus is an important tool to study changes in the language and society. The article draws attention to the problem of diachronic balance instability of various text genres in the corpus. The instability is shown not to correspond the publishing policy of the country. The method for detecting the imbalance is proposed. The Russian subcorpus is analyzed in detail to shows the presence of serious balance instabilities in the 21st century. We analyze the frequency dynamics of small groups of words, as well as make extended statistical research of the corpus. We calculate such parameters of word frequency dynamics as autocorrelations, linear trends, mathematical expectations and variances. We also apply the methods of singular spectral analysis and moving average to find unnatural peaks of the average “noise”. Our conclusion is unreasonably large amount of fiction in the Russian subcorpus in some specific years. Similar effects are also found in the English subcorpus. These facts seem to be significant because the most often studied objects of psychological and sociological research are the high-frequency common words widespread in fiction. An example is given to demonstrate that ignoring this factor can result in erroneous conclusions.