<p>In this paper, we propose semiparametric efficient estimators of genetic relatedness between two traits in a model-free framework. Most existing methods require specifying certain parametric models involving the traits and genetic variants. However, the bias due to model misspecification may yield misleading statistical results. Moreover, the semiparametric efficient bounds for estimators of genetic relatedness are still lacking. In this paper, we develop semiparametric efficient estimators with machine learning methods and construct valid confidence intervals for two important measures of genetic relatedness: genetic covariance and genetic correlation, allowing both continuous and discrete responses. Based on the derived efficient influence functions of genetic relatedness, we propose a consistent estimator of the genetic covariance as long as one of the genetic values is consistently estimated. The data of two traits may be collected from the same group or different groups of individuals. To validate our approach, we conduct various numerical studies to illustrate the introduced estimation procedures. Additionally, we apply our proposed methodologies to analyze data from the Carworth Farms White mice genome-wide association study.</p>

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

Semiparametric efficient estimation of genetic relatedness with machine learning methods

  • Xu Guo,
  • Hongwei Shi,
  • Weichao Yang,
  • Yiyuan Qian,
  • Niwen Zhou

摘要

In this paper, we propose semiparametric efficient estimators of genetic relatedness between two traits in a model-free framework. Most existing methods require specifying certain parametric models involving the traits and genetic variants. However, the bias due to model misspecification may yield misleading statistical results. Moreover, the semiparametric efficient bounds for estimators of genetic relatedness are still lacking. In this paper, we develop semiparametric efficient estimators with machine learning methods and construct valid confidence intervals for two important measures of genetic relatedness: genetic covariance and genetic correlation, allowing both continuous and discrete responses. Based on the derived efficient influence functions of genetic relatedness, we propose a consistent estimator of the genetic covariance as long as one of the genetic values is consistently estimated. The data of two traits may be collected from the same group or different groups of individuals. To validate our approach, we conduct various numerical studies to illustrate the introduced estimation procedures. Additionally, we apply our proposed methodologies to analyze data from the Carworth Farms White mice genome-wide association study.