66 research outputs found
Discerning the ancestry of European Americans in genetic association studies
European Americans are often treated as a homogeneous group, but in fact form a structured population due to historical immigration of diverse source populations. Discerning the ancestry of European Americans genotyped in association studies is important in order to prevent false-positive or false-negative associations due to population stratification and to identify genetic variants whose contribution to disease risk differs across European ancestries. Here, we investigate empirical patterns of population structure in European Americans, analyzing 4,198 samples from four genome-wide association studies to show that components roughly corresponding to northwest European, southeast European, and Ashkenazi Jewish ancestry are the main sources of European American population structure. Building on this insight, we constructed a panel of 300 validated markers that are highly informative for distinguishing these ancestries. We demonstrate that this panel of markers can be used to correct for stratification in association studies that do not generate dense genotype data
Recommended from our members
Single-Cell RNA Sequencing of Microglia throughout the Mouse Lifespan and in the Injured Brain Reveals Complex Cell-State Changes.
Microglia, the resident immune cells of the brain, rapidly change states in response to their environment, but we lack molecular and functional signatures of different microglial populations. Here, we analyzed the RNA expression patterns of more than 76,000 individual microglia in mice during development, in old age, and after brain injury. Our analysis uncovered at least nine transcriptionally distinct microglial states, which expressed unique sets of genes and were localized in the brain using specific markers. The greatest microglial heterogeneity was found at young ages; however, several states-including chemokine-enriched inflammatory microglia-persisted throughout the lifespan or increased in the aged brain. Multiple reactive microglial subtypes were also found following demyelinating injury in mice, at least one of which was also found in human multiple sclerosis lesions. These distinct microglia signatures can be used to better understand microglia function and to identify and manipulate specific subpopulations in health and disease
Integrating sequence and array data to create an improved 1000 Genomes Project haplotype reference panel
A major use of the 1000 Genomes Project (1000GP) data is genotype imputation in genome-wide association studies (GWAS). Here we develop a method to estimate haplotypes from low-coverage sequencing data that can take advantage of single-nucleotide polymorphism (SNP) microarray genotypes on the same samples. First the SNP array data are phased to build a backbone (or \u27scaffold\u27) of haplotypes across each chromosome. We then phase the sequence data \u27onto\u27 this haplotype scaffold. This approach can take advantage of relatedness between sequenced and non-sequenced samples to improve accuracy. We use this method to create a new 1000GP haplotype reference set for use by the human genetic community. Using a set of validation genotypes at SNP and bi-allelic indels we show that these haplotypes have lower genotype discordance and improved imputation performance into downstream GWAS samples, especially at low-frequency variants. © 2014 Macmillan Publishers Limited. All rights reserved
Recommended from our members
Mapping Copy Number Variation by Population Scale Genome Sequencing
Genomic structural variants (SVs) are abundant in humans, differing from other forms of variation in extent, origin and functional impact. Despite progress in SV characterization, the nucleotide resolution architecture of most SVs remains unknown. We constructed a map of unbalanced SVs (that is, copy number variants) based on whole genome DNA sequencing data from 185 human genomes, integrating evidence from complementary SV discovery approaches with extensive experimental validations. Our map encompassed 22,025 deletions and 6,000 additional SVs, including insertions and tandem duplications. Most SVs (53%) were mapped to nucleotide resolution, which facilitated analysing their origin and functional impact. We examined numerous whole and partial gene deletions with a genotyping approach and observed a depletion of gene disruptions amongst high frequency deletions. Furthermore, we observed differences in the size spectra of SVs originating from distinct formation mechanisms, and constructed a map of SV hotspots formed by common mechanisms. Our analytical framework and SV map serves as a resource for sequencing-based association studies.Organismic and Evolutionary Biolog
Recommended from our members
Characterization of Bipolar Disorder Patient-Specific Induced Pluripotent Stem Cells from a Family Reveals Neurodevelopmental and mRNA Expression Abnormalities
Bipolar disorder (BD) is a common neuropsychiatric disorder characterized by chronic recurrent episodes of depression and mania. Despite evidence for high heritability of BD, little is known about its underlying pathophysiology. To develop new tools for investigating the molecular and cellular basis of BD we applied a family-based paradigm to derive and characterize a set of 12 induced pluripotent stem cell (iPSC) lines from a quartet consisting of two BD-affected brothers and their two unaffected parents. Initially, no significant phenotypic differences were observed between iPSCs derived from the different family members. However, upon directed neural differentiation we observed that CXCR4 (CXC chemokine receptor-4) expressing central nervous system (CNS) neural progenitor cells (NPCs) from both BD patients compared to their unaffected parents exhibited multiple phenotypic differences at the level of neurogenesis and expression of genes critical for neuroplasticity, including WNT pathway components and ion channel subunits. Treatment of the CXCR4+ NPCs with a pharmacological inhibitor of glycogen synthase kinase 3 (GSK3), a known regulator of WNT signaling, was found to rescue a progenitor proliferation deficit in the BD-patient NPCs. Taken together, these studies provide new cellular tools for dissecting the pathophysiology of BD and evidence for dysregulation of key pathways involved in neurodevelopment and neuroplasticity. Future generation of additional iPSCs following a family-based paradigm for modeling complex neuropsychiatric disorders in conjunction with in-depth phenotyping holds promise for providing insights into the pathophysiological substrates of BD and is likely to inform the development of targeted therapeutics for its treatment and ideally prevention
Integrating sequence and array data to create an improved 1000 Genomes Project haplotype reference panel
A major use of the 1000 Genomes Project (1000GP) data is genotype imputation in genome-wide association studies (GWAS). Here we develop a method to estimate haplotypes from low-coverage sequencing data that can take advantage of single-nucleotide polymorphism (SNP) microarray genotypes on the same samples. First the SNP array data are phased to build a backbone (or 'scaffold') of haplotypes across each chromosome. We then phase the sequence data 'onto' this haplotype scaffold. This approach can take advantage of relatedness between sequenced and non-sequenced samples to improve accuracy. We use this method to create a new 1000GP haplotype reference set for use by the human genetic community. Using a set of validation genotypes at SNP and bi-allelic indels we show that these haplotypes have lower genotype discordance and improved imputation performance into downstream GWAS samples, especially at low-frequency variants. © 2014 Macmillan Publishers Limited. All rights reserved
Association analyses of 249,796 individuals reveal 18 new loci associated with body mass index
Obesity is globally prevalent and highly heritable, but the underlying genetic factors remain largely elusive. To identify genetic loci for obesity-susceptibility, we examined associations between body mass index (BMI) and ~2.8 million SNPs in up to 123,865 individuals, with targeted follow-up of 42 SNPs in up to 125,931 additional individuals. We confirmed 14 known obesity-susceptibility loci and identified 18 new loci associated with BMI (P<5×10−8), one of which includes a copy number variant near GPRC5B. Some loci (MC4R, POMC, SH2B1, BDNF) map near key hypothalamic regulators of energy balance, and one is near GIPR, an incretin receptor. Furthermore, genes in other newly-associated loci may provide novel insights into human body weight regulation
Genome-wide association identifies nine common variants associated with fasting proinsulin levels and provides new insights into the pathophysiology of type 2 diabetes.
OBJECTIVE: Proinsulin is a precursor of mature insulin and C-peptide. Higher circulating proinsulin levels are associated with impaired β-cell function, raised glucose levels, insulin resistance, and type 2 diabetes (T2D). Studies of the insulin processing pathway could provide new insights about T2D pathophysiology. RESEARCH DESIGN AND METHODS: We have conducted a meta-analysis of genome-wide association tests of ∼2.5 million genotyped or imputed single nucleotide polymorphisms (SNPs) and fasting proinsulin levels in 10,701 nondiabetic adults of European ancestry, with follow-up of 23 loci in up to 16,378 individuals, using additive genetic models adjusted for age, sex, fasting insulin, and study-specific covariates. RESULTS: Nine SNPs at eight loci were associated with proinsulin levels (P < 5 × 10(-8)). Two loci (LARP6 and SGSM2) have not been previously related to metabolic traits, one (MADD) has been associated with fasting glucose, one (PCSK1) has been implicated in obesity, and four (TCF7L2, SLC30A8, VPS13C/C2CD4A/B, and ARAP1, formerly CENTD2) increase T2D risk. The proinsulin-raising allele of ARAP1 was associated with a lower fasting glucose (P = 1.7 × 10(-4)), improved β-cell function (P = 1.1 × 10(-5)), and lower risk of T2D (odds ratio 0.88; P = 7.8 × 10(-6)). Notably, PCSK1 encodes the protein prohormone convertase 1/3, the first enzyme in the insulin processing pathway. A genotype score composed of the nine proinsulin-raising alleles was not associated with coronary disease in two large case-control datasets. CONCLUSIONS: We have identified nine genetic variants associated with fasting proinsulin. Our findings illuminate the biology underlying glucose homeostasis and T2D development in humans and argue against a direct role of proinsulin in coronary artery disease pathogenesis
Contributions of common genetic variants to risk of schizophrenia among individuals of African and Latino ancestry
Schizophrenia is a common, chronic and debilitating neuropsychiatric syndrome affecting tens of millions of individuals worldwide. While rare genetic variants play a role in the etiology of schizophrenia, most of the currently explained liability is within common variation, suggesting that variation predating the human diaspora out of Africa harbors a large fraction of the common variant attributable heritability. However, common variant association studies in schizophrenia have concentrated mainly on cohorts of European descent. We describe genome-wide association studies of 6152 cases and 3918 controls of admixed African ancestry, and of 1234 cases and 3090 controls of Latino ancestry, representing the largest such study in these populations to date. Combining results from the samples with African ancestry with summary statistics from the Psychiatric Genomics Consortium (PGC) study of schizophrenia yielded seven newly genome-wide significant loci, and we identified an additional eight loci by incorporating the results from samples with Latino ancestry. Leveraging population differences in patterns of linkage disequilibrium, we achieve improved fine-mapping resolution at 22 previously reported and 4 newly significant loci. Polygenic risk score profiling revealed improved prediction based on trans-ancestry meta-analysis results for admixed African (Nagelkerke’s R2 = 0.032; liability R2 = 0.017; P < 10−52), Latino (Nagelkerke’s R2 = 0.089; liability R2 = 0.021; P < 10−58), and European individuals (Nagelkerke’s R2 = 0.089; liability R2 = 0.037; P < 10−113), further highlighting the advantages of incorporating data from diverse human populations
Integrating sequence and array data to create an improved 1000 Genomes Project haplotype reference panel
A major use of the 1000 Genomes Project (1000GP) data is genotype imputation in genome-wide association studies (GWAS). Here we develop a method to estimate haplotypes from low coverage sequencing data that can take advantage of SNP microarray genotypes on the same samples. Firstly the SNP array data are phased in order to build a backbone (or ’scaffold’) of haplotypes across each chromosome. We then phase the sequence data ‘onto’ this haplotype scaffold. This approach can take advantage of relatedness between sequenced and non-sequenced samples to improve accuracy. We use this method to create a new 1000GP haplotype reference set for use by the human genetic community. Using a set of validation genotypes at SNP and biallelic indels we show that these haplotypes have lower genotype discordance and improved imputation performance into downstream GWAS samples, especially at low frequency variants
- …