The Benford law is a remarkable and interesting phenomenon. It tells that the significand of certain data sets is not uniformly distributed. Especially, the probability that the first nonzero digit equals One is approximately 0.3. It turns out that the significand has a logarithmic distribution. The Benford random variables have interesting properties, such as scale invariance, base invariance, and sum invariance. Moreover, mixtures of random variables are often also Benford distributed. Several tests exist to test whether or not a data set is Benford distributed. Among them are Pearson’s \(\chi ^2\) -test, the Kolmogorov-Smirnov test, and the MAD test. Usually, these tests are applied to the first significant digit only. However, similar tests may be constructed based on the second significant digit or other significant digits. Recently, the authors proposed further tests based on the sum- and scale-invariance properties, respectively. These tests are based upon the underlying significands. The main idea of those tests is presented. Benford’s law can be applied to detect fraud or nonconforming entries in data sets. We illustrate the various tests on two theoretical and two empirical data sets. Moreover, we discuss types of data fraud and design a scenario of a specific kind of manipulation. It allows studying the power of those tests.

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

On Some Properties and Testing of Benford’s Law

  • Wolfgang Kössler,
  • Hans-J. Lenz,
  • Xing D. Wang

摘要

The Benford law is a remarkable and interesting phenomenon. It tells that the significand of certain data sets is not uniformly distributed. Especially, the probability that the first nonzero digit equals One is approximately 0.3. It turns out that the significand has a logarithmic distribution. The Benford random variables have interesting properties, such as scale invariance, base invariance, and sum invariance. Moreover, mixtures of random variables are often also Benford distributed. Several tests exist to test whether or not a data set is Benford distributed. Among them are Pearson’s \(\chi ^2\) -test, the Kolmogorov-Smirnov test, and the MAD test. Usually, these tests are applied to the first significant digit only. However, similar tests may be constructed based on the second significant digit or other significant digits. Recently, the authors proposed further tests based on the sum- and scale-invariance properties, respectively. These tests are based upon the underlying significands. The main idea of those tests is presented. Benford’s law can be applied to detect fraud or nonconforming entries in data sets. We illustrate the various tests on two theoretical and two empirical data sets. Moreover, we discuss types of data fraud and design a scenario of a specific kind of manipulation. It allows studying the power of those tests.