<p>In many modern applications, analysts observe only aggregated versions of discrete data, such as intervals or fixed-bin frequency summaries built from latent counts. Within the Symbolic Data Analysis framework, these aggregates are <i>symbols</i> generated by a known mapping of unobserved micro-data. Building on the general symbolic-likelihood perspective for aggregated data, this paper develops a likelihood-based treatment tailored to symbolic data generated from latent discrete count models. We construct the <i>symbolic likelihood</i> by summing the latent model over all configurations compatible with each observed symbol, and we derive a score identity showing that the symbolic score equals the conditional expectation of the complete-data score given the symbol. For discrete exponential-family models, this leads to simple estimating equations and EM-type updates that mirror their full-data counterparts; in the Poisson fixed-bin frequency case, the resulting fixed-point iteration is particularly tractable. A simulation study quantifies the efficiency loss induced by different symbol designs, showing that fixed-bin frequency symbolic maximum likelihood estimators recover most of the information in the full data, while min–max intervals and midpoint heuristics can perform noticeably worse. The methodology is illustrated with NBA free-throw attempt data, where only season-level grouped-count summaries of game-level counts are used. Symbolic Poisson and Negative Binomial models are fitted and compared, and goodness-of-fit is assessed entirely in the symbol space via Pearson-type statistics and parametric bootstrap.</p>

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

Symbolic likelihood inference for discrete aggregated count data: an NBA application

  • Abdolnasser Sadeghkhani

摘要

In many modern applications, analysts observe only aggregated versions of discrete data, such as intervals or fixed-bin frequency summaries built from latent counts. Within the Symbolic Data Analysis framework, these aggregates are symbols generated by a known mapping of unobserved micro-data. Building on the general symbolic-likelihood perspective for aggregated data, this paper develops a likelihood-based treatment tailored to symbolic data generated from latent discrete count models. We construct the symbolic likelihood by summing the latent model over all configurations compatible with each observed symbol, and we derive a score identity showing that the symbolic score equals the conditional expectation of the complete-data score given the symbol. For discrete exponential-family models, this leads to simple estimating equations and EM-type updates that mirror their full-data counterparts; in the Poisson fixed-bin frequency case, the resulting fixed-point iteration is particularly tractable. A simulation study quantifies the efficiency loss induced by different symbol designs, showing that fixed-bin frequency symbolic maximum likelihood estimators recover most of the information in the full data, while min–max intervals and midpoint heuristics can perform noticeably worse. The methodology is illustrated with NBA free-throw attempt data, where only season-level grouped-count summaries of game-level counts are used. Symbolic Poisson and Negative Binomial models are fitted and compared, and goodness-of-fit is assessed entirely in the symbol space via Pearson-type statistics and parametric bootstrap.