Search CORE

255 research outputs found

Recommended from our members

Improving the Power of GWAS and Avoiding Confounding from Population Stratification with PC-Select

Author: Berger Bonnie
Price Alkes L.
Tucker George
Publication venue: 'Genetics Society of America'
Publication date: 13/08/2014
Field of study

Using a reduced subset of SNPs in a linear mixed model can improve power for genome-wide association studies, yet this can result in insufficient correction for population stratification. We propose a hybrid approach using principal components that does not inflate statistics in the presence of population stratification and improves power over standard linear mixed models

Harvard University - DASH

Application of Ancestry Informative Markers to Association Studies in European Americans

Author: Price Alkes L
Seldin Michael F
Publication venue: Public Library of Science
Publication date: 01/01/2008
Field of study

Directory of Open Access Journals

PubMed Central

Population Structure and Eigenanalysis

Author: Patterson Nick
Price Alkes L
Reich David
Publication venue: Public Library of Science
Publication date: 01/01/2006
Field of study

Current methods for inferring population structure from genetic data do not provide formal significance tests for population differentiation. We discuss an approach to studying population structure (principal components analysis) that was first applied to genetic data by Cavalli-Sforza and colleagues. We place the method on a solid statistical footing, using results from modern statistics to develop formal significance tests. We also uncover a general “phase change” phenomenon about the ability to detect structure in genetic data, which emerges from the statistical theory we use, and has an important implication for the ability to discover structure in genetic data: for a fixed but large dataset size, divergence between two populations (as measured, for example, by a statistic like F(ST)) below a threshold is essentially undetectable, but a little above threshold, detection will be easy. This means that we can predict the dataset size needed to detect structure

CiteSeerX

Public Library of Science (PLOS)

Harvard University - DASH

Directory of Open Access Journals

PubMed Central

Explicit Modeling of Ancestry Improves Polygenic Risk Scores and BLUP Prediction

Author: Chen Chia-Yen
Han Jiali
Hunter David J.
Kraft Peter
Price Alkes L.
Publication venue: 'Wiley'
Publication date: 01/09/2015
Field of study

Polygenic prediction using genome-wide SNPs can provide high prediction accuracy for complex traits. Here, we investigate the question of how to account for genetic ancestry when conducting polygenic prediction. We show that the accuracy of polygenic prediction in structured populations may be partly due to genetic ancestry. However, we hypothesized that explicitly modeling ancestry could improve polygenic prediction accuracy. We analyzed three GWAS of hair color (HC), tanning ability (TA), and basal cell carcinoma (BCC) in European Americans (sample size from 7,440 to 9,822) and considered two widely used polygenic prediction approaches: polygenic risk scores (PRSs) and best linear unbiased prediction (BLUP). We compared polygenic prediction without correction for ancestry to polygenic prediction with ancestry as a separate component in the model. In 10-fold cross-validation using the PRS approach, the R(2) for HC increased by 66% (0.0456-0.0755; P < 10(-16)), the R(2) for TA increased by 123% (0.0154 to 0.0344; P < 10(-16)), and the liability-scale R(2) for BCC increased by 68% (0.0138-0.0232; P < 10(-16)) when explicitly modeling ancestry, which prevents ancestry effects from entering into each SNP effect and being overweighted. Surprisingly, explicitly modeling ancestry produces a similar improvement when using the BLUP approach, which fits all SNPs simultaneously in a single variance component and causes ancestry to be underweighted. We validate our findings via simulations, which show that the differences in prediction accuracy will increase in magnitude as sample sizes increase. In summary, our results show that explicitly modeling ancestry can be important in both PRS and BLUP prediction

IUPUIScholarWorks

PubMed Central

Integrating Functional Data to Prioritize Causal Variants in Statistical Fine-Mapping Studies

Author: Eskin Eleazar
Hormozdiari Farhad
Kichaev Gleb
Kraft Peter
Lindstrom Sara
Pasaniuc Bogdan
Price Alkes L.
Yang Wen-Yun
Publication venue: 'Public Library of Science (PLoS)'
Publication date: 01/10/2014
Field of study

Standard statistical approaches for prioritization of variants for functional testing in fine-mapping studies either use marginal association statistics or estimate posterior probabilities for variants to be causal under simplifying assumptions. Here, we present a probabilistic framework that integrates association strength with functional genomic annotation data to improve accuracy in selecting plausible causal variants for functional validation. A key feature of our approach is that it empirically estimates the contribution of each functional annotation to the trait of interest directly from summary association statistics while allowing for multiple causal variants at any risk locus. We devise efficient algorithms that estimate the parameters of our model across all risk loci to further increase performance. Using simulations starting from the 1000 Genomes data, we find that our framework consistently outperforms the current state-of-the-art fine-mapping methods, reducing the number of variants that need to be selected to capture 90% of the causal variants from an average of 13.3 to 10.4 SNPs per locus (as compared to the next-best performing strategy). Furthermore, we introduce a cost-to-benefit optimization framework for determining the number of variants to be followed up in functional assays and assess its performance using real and simulation data. We validate our findings using a large scale meta-analysis of four blood lipids traits and find that the relative probability for causality is increased for variants in exons and transcription start sites and decreased in repressed genomic regions at the risk loci of these traits. Using these highly predictive, trait-specific functional annotations, we estimate causality probabilities across all traits and variants, reducing the size of the 90% confidence set from an average of 17.5 to 13.5 variants per locus in this data

Crossref

Harvard University - DASH

Directory of Open Access Journals

PubMed Central

eScholarship - University of California

Recommended from our members

Fast and accurate long-range phasing in a UK Biobank cohort

Author: Loh Po-Ru
Palamara Pier Francesco
Price Alkes L
Publication venue: 'Springer Science and Business Media LLC'
Publication date: 03/01/2017
Field of study

Recent work has leveraged the extensive genotyping of the Icelandic population to perform long-range phasing (LRP), enabling accurate imputation and association analysis of rare variants in target samples typed on genotyping arrays. Here, we develop a fast and accurate LRP method, Eagle, that extends this paradigm to populations with much smaller proportions of genotyped samples by harnessing long (>4cM) identical-by-descent (IBD) tracts shared among distantly related individuals. We applied Eagle to N≈150,000 samples (0.2% of the British population) from the UK Biobank, and we determined that it is 1–2 orders of magnitude faster than existing methods while achieving similar or better phasing accuracy (switch error rate ≈0.3%, corresponding to perfect phase in a majority of 10Mb segments). We also observed that when used within an imputation pipeline, Eagle pre-phasing improved downstream imputation accuracy compared to pre-phasing in batches using existing methods (as necessary to achieve comparable computational cost)

Harvard University - DASH

Progress and promise in understanding the genetic basis of common diseases

Author: Donnelly Peter
Price Alkes L.
Spencer Chris C. A.
Publication venue: 'The Royal Society'
Publication date: 01/01/2015
Field of study

Susceptibility to common human diseases is influenced by both genetic and environmental factors. The explosive growth of genetic data, and the knowledge that it is generating, are transforming our biological understanding of these diseases. In this review, we describe the technological and analytical advances that have enabled genome-wide association studies to be successful in identifying a large number of genetic variants robustly associated with common disease. We examine the biological insights that these genetic associations are beginning to produce, from functional mechanisms involving individual genes to biological pathways linking associated genes, and the identification of functional annotations, some of which are cell-type-specific, enriched in disease associations. Although most efforts have focused on identifying and interpreting genetic variants that are irrefutably associated with disease, it is increasingly clear that—even at large sample sizes—these represent only the tip of the iceberg of genetic signal, motivating polygenic analyses that consider the effects of genetic variants throughout the genome, including modest effects that are not individually statistically significant. As data from an increasingly large number of diseases and traits are analysed, pleiotropic effects (defined as genetic loci affecting multiple phenotypes) can help integrate our biological understanding. Looking forward, the next generation of population-scale data resources, linking genomic information with health outcomes, will lead to another step-change in our ability to understand, and treat, common diseases

Harvard University - DASH

Oxford University Research Archive

Leveraging population admixture to characterize the heritability of complex traits

Author: Kittles Rick A
Price Alkes L
Reiner Alex P
Stram Daniel
Tang Hua
Publication venue
Publication date: 01/01/2014
Field of study

Despite recent progress on estimating the heritability explained by genotyped SNPs (hg2), a large gap between hg2 and estimates of total narrow-sense heritability (h2) remains. Explanations for this gap include rare variants, or upward bias in family-based estimates of h2 due to shared environment or epistasis. We estimate h2 from unrelated individuals in admixed populations by first estimating the heritability explained by local ancestry (hγ2). We show that hγ2 = 2FSTCθ(1−θ)h2, where FSTC measures frequency differences between populations at causal loci and θ is the genome-wide ancestry proportion. Our approach is not susceptible to biases caused by epistasis or shared environment. We examined 21,497 African Americans from three cohorts, analyzing 13 phenotypes. For height and BMI, we obtained h2 estimates of 0.55 ± 0.09 and 0.23 ± 0.06, respectively, which are larger than estimates of hg2 in these and other data, but smaller than family-based estimates of h2

Carolina Digital Repository