4.7 Article

Large-scale inference of population structure in presence of missingness using PCA

Journal

BIOINFORMATICS
Volume 37, Issue 13, Pages 1868-1875

Publisher

OXFORD UNIV PRESS
DOI: 10.1093/bioinformatics/btab027

Keywords

-

Funding

  1. Lundbeck foundation [R215-2015-4174]
  2. National Natural Science Foundation of China [31900487]

Ask authors/readers for more resources

The article introduces a method EMU for inferring population structure in the presence of massive missing data, which is both fast and accurate. Testing on the Chinese Millionome Project dataset demonstrates the effectiveness of EMU in capturing population structure.
Motivation: Principal component analysis (PCA) is a commonly used tool in genetics to capture and visualize population structure. Due to technological advances in sequencing, such as the widely used non-invasive prenatal test, massive datasets of ultra-low coverage sequencing are being generated. These datasets are characterized by having a large amount of missing genotype information. Results: We present EMU, a method for inferring population structure in the presence of rampant non-random missingness. We show through simulations that several commonly used PCA methods cannot handle missing data arisen from various sources, which leads to biased results as individuals are projected into the PC space based on their amount of missingness. In terms of accuracy, EMU outperforms an existing method that also accommodates missingness while being competitively fast. We further tested EMU on around 100K individuals of the Phase 1 dataset of the Chinese Millionome Project, that were shallowly sequenced to around 0.08x. From this data we are able to capture the population structure of the Han Chinese and to reproduce previous analysis in a matter of CPU hours instead of CPU years. EMU's capability to accurately infer population structure in the presence of missingness will be of increasing importance with the rising number of large-scale genetic datasets.

Authors

I am an author on this paper
Click your name to claim this paper and add it to your profile.

Reviews

Primary Rating

4.7
Not enough ratings

Secondary Ratings

Novelty
-
Significance
-
Scientific rigor
-
Rate this paper

Recommended

No Data Available
No Data Available