Even big data does not speak for itself
摘要
This paper demonstrates that in an important class of data-integration problems even Big Data cannot speak for itself: marginal probabilities provided by large and unbiased datasets do not determine a unique joint probability distribution. Big Data must instead be accompanied by further reasoning methodologies for handling unavoidable uncertainties. I assess how Bayesian approaches handle this underdetermination and argue that they face practical challenges when dealing with Big Data that does not speak for itself.