<p>This paper addresses the challenge of variable selection and statistical inference in multiple compositional data, particularly in the context of multi-site human microbiome studies. We propose a novel extension of the group lasso framework to handle compositional covariates through the multiple compositional regression model. Our approach, which incorporates both compositional constraints and a debiasing procedure, enables asymptotically unbiased and normally distributed estimates for grouped coefficients. We develop statistical significance tests based on asymptotic chi-squared statistics for grouped variables and provide rigorous theoretical guarantees on estimation error bounds. Through extensive simulation studies, we demonstrate that the proposed method outperforms traditional approaches, especially in scenarios with weak signals, by achieving higher positive predictive values. Finally, we apply the method to a human microbiome dataset to identify bacterial taxa associated with cholesterol levels, showcasing the practical utility of our approach.</p>

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

Debiased group lasso for multiple compositional data

  • Sujin Lee,
  • Sungkyu Jung

摘要

This paper addresses the challenge of variable selection and statistical inference in multiple compositional data, particularly in the context of multi-site human microbiome studies. We propose a novel extension of the group lasso framework to handle compositional covariates through the multiple compositional regression model. Our approach, which incorporates both compositional constraints and a debiasing procedure, enables asymptotically unbiased and normally distributed estimates for grouped coefficients. We develop statistical significance tests based on asymptotic chi-squared statistics for grouped variables and provide rigorous theoretical guarantees on estimation error bounds. Through extensive simulation studies, we demonstrate that the proposed method outperforms traditional approaches, especially in scenarios with weak signals, by achieving higher positive predictive values. Finally, we apply the method to a human microbiome dataset to identify bacterial taxa associated with cholesterol levels, showcasing the practical utility of our approach.