<p>Accurately predicting the effect of missense variants is important in discovering disease risk genes and clinical genetic diagnostics. Commonly used computational methods predict pathogenicity, which does not capture the quantitative impact on fitness in humans. We develop a method, MisFit, to estimate missense fitness effect using a graphical model. MisFit jointly models the effect at a molecular level (<InlineEquation ID="IEq1"> <InlineMediaObject> <ImageObject Color="BlackWhite" FileRef="41467_2025_59937_Article_IEq1.gif" Format="GIF" Height="14" Rendition="HTML" Resolution="72" Type="Linedraw" Width="13" /> </InlineMediaObject> <EquationSource Format="TEX">\(d\)</EquationSource> <EquationSource Format="MATHML"><math> <mi>d</mi> </math></EquationSource> </InlineEquation>) and a population level (selection coefficient, <InlineEquation ID="IEq2"> <InlineMediaObject> <ImageObject Color="BlackWhite" FileRef="41467_2025_59937_Article_IEq2.gif" Format="GIF" Height="10" Rendition="HTML" Resolution="72" Type="Linedraw" Width="12" /> </InlineMediaObject> <EquationSource Format="TEX">\(s\)</EquationSource> <EquationSource Format="MATHML"><math> <mi>s</mi> </math></EquationSource> </InlineEquation>), assuming that in the same gene, missense variants with similar <InlineEquation ID="IEq3"> <InlineMediaObject> <ImageObject Color="BlackWhite" FileRef="41467_2025_59937_Article_IEq1.gif" Format="GIF" Height="14" Rendition="HTML" Resolution="72" Type="Linedraw" Width="13" /> </InlineMediaObject> <EquationSource Format="TEX">\(d\)</EquationSource> <EquationSource Format="MATHML"><math> <mi>d</mi> </math></EquationSource> </InlineEquation> have similar <InlineEquation ID="IEq4"> <InlineMediaObject> <ImageObject Color="BlackWhite" FileRef="41467_2025_59937_Article_IEq2.gif" Format="GIF" Height="10" Rendition="HTML" Resolution="72" Type="Linedraw" Width="12" /> </InlineMediaObject> <EquationSource Format="TEX">\(s\)</EquationSource> <EquationSource Format="MATHML"><math> <mi>s</mi> </math></EquationSource> </InlineEquation>. We train it by maximizing probability of observed allele counts in 236,017 individuals of European ancestry. We show that <InlineEquation ID="IEq5"> <InlineMediaObject> <ImageObject Color="BlackWhite" FileRef="41467_2025_59937_Article_IEq2.gif" Format="GIF" Height="10" Rendition="HTML" Resolution="72" Type="Linedraw" Width="12" /> </InlineMediaObject> <EquationSource Format="TEX">\(s\)</EquationSource> <EquationSource Format="MATHML"><math> <mi>s</mi> </math></EquationSource> </InlineEquation> is informative in predicting allele frequency across ancestries and consistent with the fraction of de novo mutations in sites under strong selection. Further, <InlineEquation ID="IEq6"> <InlineMediaObject> <ImageObject Color="BlackWhite" FileRef="41467_2025_59937_Article_IEq2.gif" Format="GIF" Height="10" Rendition="HTML" Resolution="72" Type="Linedraw" Width="12" /> </InlineMediaObject> <EquationSource Format="TEX">\(s\)</EquationSource> <EquationSource Format="MATHML"><math> <mi>s</mi> </math></EquationSource> </InlineEquation> outperforms previous methods in prioritizing de novo missense variants in individuals with neurodevelopmental disorders. In conclusion, MisFit accurately predicts <InlineEquation ID="IEq7"> <InlineMediaObject> <ImageObject Color="BlackWhite" FileRef="41467_2025_59937_Article_IEq2.gif" Format="GIF" Height="10" Rendition="HTML" Resolution="72" Type="Linedraw" Width="12" /> </InlineMediaObject> <EquationSource Format="TEX">\(s\)</EquationSource> <EquationSource Format="MATHML"><math> <mi>s</mi> </math></EquationSource> </InlineEquation> and yields new insights from genomic data.</p>

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

A probabilistic graphical model for estimating selection coefficients of nonsynonymous variants from human population sequence data

  • Yige Zhao,
  • Tian Lan,
  • Guojie Zhong,
  • Jake Hagen,
  • Hongbing Pan,
  • Wendy K. Chung,
  • Yufeng Shen

摘要

Accurately predicting the effect of missense variants is important in discovering disease risk genes and clinical genetic diagnostics. Commonly used computational methods predict pathogenicity, which does not capture the quantitative impact on fitness in humans. We develop a method, MisFit, to estimate missense fitness effect using a graphical model. MisFit jointly models the effect at a molecular level ( \(d\) d ) and a population level (selection coefficient, \(s\) s ), assuming that in the same gene, missense variants with similar \(d\) d have similar \(s\) s . We train it by maximizing probability of observed allele counts in 236,017 individuals of European ancestry. We show that \(s\) s is informative in predicting allele frequency across ancestries and consistent with the fraction of de novo mutations in sites under strong selection. Further, \(s\) s outperforms previous methods in prioritizing de novo missense variants in individuals with neurodevelopmental disorders. In conclusion, MisFit accurately predicts \(s\) s and yields new insights from genomic data.