Braun Rosemary, Rowe William, Schaefer Carl, Zhang Jinghui, Buetow Kenneth
Laboratory of Population Genetics, National Cancer Institute, National Institutes of Health, Bethesda, Maryland, United States of America.
PLoS Genet. 2009 Oct;5(10):e1000668. doi: 10.1371/journal.pgen.1000668. Epub 2009 Oct 2.
Recent publications have described and applied a novel metric that quantifies the genetic distance of an individual with respect to two population samples, and have suggested that the metric makes it possible to infer the presence of an individual of known genotype in a sample for which only the marginal allele frequencies are known. However, the assumptions, limitations, and utility of this metric remained incompletely characterized. Here we present empirical tests of the method using publicly accessible genotypes, as well as analytical investigations of the method's strengths and limitations. The results reveal that the null distribution is sensitive to the underlying assumptions, making it difficult to accurately calibrate thresholds for classifying an individual as a member of the population samples. As a result, the false-positive rates obtained in practice are considerably higher than previously believed. However, despite the metric's inadequacies for identifying the presence of an individual in a sample, our results suggest potential avenues for future research on tuning this method to problems of ancestry inference or disease prediction. By revealing both the strengths and limitations of the proposed method, we hope to elucidate situations in which this distance metric may be used in an appropriate manner. We also discuss the implications of our findings in forensics applications and in the protection of GWAS participant privacy.
近期的出版物描述并应用了一种新的度量标准,该标准可量化个体相对于两个群体样本的遗传距离,并表明该度量标准能够在仅知道边际等位基因频率的样本中推断出已知基因型个体的存在。然而,该度量标准的假设、局限性和实用性仍未得到充分描述。在此,我们使用公开可用的基因型对该方法进行实证检验,并对该方法的优势和局限性进行分析研究。结果表明,零分布对潜在假设敏感,这使得难以准确校准将个体分类为群体样本成员的阈值。因此,实际获得的假阳性率远高于先前的认知。然而,尽管该度量标准在识别样本中个体的存在方面存在不足,但我们的结果为未来将该方法调整用于血统推断或疾病预测问题的研究提供了潜在途径。通过揭示所提出方法的优势和局限性,我们希望阐明可以适当使用这种距离度量标准的情况。我们还讨论了我们的研究结果在法医学应用以及保护全基因组关联研究(GWAS)参与者隐私方面的意义。