813 research outputs found

    Epistatic Module Detection for Case-Control Studies: A Bayesian Model with a Gibbs Sampling Strategy

    Get PDF
    The detection of epistatic interactive effects of multiple genetic variants on the susceptibility of human complex diseases is a great challenge in genome-wide association studies (GWAS). Although methods have been proposed to identify such interactions, the lack of an explicit definition of epistatic effects, together with computational difficulties, makes the development of new methods indispensable. In this paper, we introduce epistatic modules to describe epistatic interactive effects of multiple loci on diseases. On the basis of this notion, we put forward a Bayesian marker partition model to explain observed case-control data, and we develop a Gibbs sampling strategy to facilitate the detection of epistatic modules. Comparisons of the proposed approach with three existing methods on seven simulated disease models demonstrate the superior performance of our approach. When applied to a genome-wide case-control data set for Age-related Macular Degeneration (AMD), the proposed approach successfully identifies two known susceptible loci and suggests that a combination of two other loci—one in the gene SGCD and the other in SCAPER—is associated with the disease. Further functional analysis supports the speculation that the interaction of these two genetic variants may be responsible for the susceptibility of AMD. When applied to a genome-wide case-control data set for Parkinson's disease, the proposed method identifies seven suspicious loci that may contribute independently to the disease

    Reliable confidence intervals in quantitative genetics: narrow-sense heritability

    Get PDF
    Many quantitative genetic statistics are functions of variance components, for which a large number of replicates is needed for precise estimates and reliable measures of uncertainty, on which sound interpretation depends. Moreover, in large experiments the deaths of some individuals can occur, so methods for analysing such data need to be robust to missing values. We show how confidence intervals for narrow-sense heritability can be calculated in a nested full-sib/half-sib breeding design (males crossed with several females) in the presence of missing values. Simulations indicate that the method provides accurate results, and that estimator uncertainty is lowest for sampling designs with many males relative to the number of females per male, and with more females per male than progenies per female. Missing data generally had little influence on estimator accuracy, thus suggesting that the overall number of observations should be increased even if this results in unbalanced data. We also suggest the use of parametrically simulated data for prior investigation of the accuracy of planned experiments. Together with the proposed confidence intervals an informed decision on the optimal sampling design is possible, which allows efficient allocation of resource

    Temporal and genomic analysis of additive genetic variance in breeding programmes

    Get PDF
    Genetic variance is a central parameter in quantitative genetics and breeding. Assessing changes in genetic variance over time as well as the genome is therefore of high interest. Here, we extend a previously proposed framework for temporal analysis of genetic variance using the pedigree-based model, to a new framework for temporal and genomic analysis of genetic variance using marker-based models. To this end, we describe the theory of partitioning genetic variance into genic variance and within-chromosome and between-chromosome linkage-disequilibrium, and how to estimate these variance components from a marker-based model fitted to observed phenotype and marker data. The new framework involves three steps: (i) fitting a marker-based model to data, (ii) sampling realisations of marker effects from the fitted model and for each sample calculating realisations of genetic values and (iii) calculating the variance of sampled genetic values by time and genome partitions. Analysing time partitions indicates breeding programme sustainability, while analysing genome partitions indicates contributions from chromosomes and chromosome pairs and linkage-disequilibrium. We demonstrate the framework with a simulated breeding programme involving a complex trait. Results show good concordance between simulated and estimated variances, provided that the fitted model is capturing genetic complexity of a trait. We observe a reduction of genetic variance due to selection and drift changing allele frequencies, and due to selection inducing negative linkage-disequilibrium

    Contrasting multi-site genotypic distributions among discordant quantitative phenotypes: the APOA1/C3/A4/A5 gene cluster and cardiovascular disease risk factors

    Full text link
    Most tests of association between DNA sequence variation and quantitative phenotypes in samples of randomly chosen individuals rely on specification of genotypic strata followed by comparison of phenotypes across these strata. This strategy often succeeds when phenotypic differences are caused by one or two single nucleotide polymorphisms (SNPs) among the surveyed markers. However, when multiple-SNP haplotypes account for observed phenotypic variation, identification of the best partitioning requires examination of an inordinate number of SNP combinations. An alternative approach is to rank individuals by their phenotypic measures and ask whether attributes of the genotypic variation show a non-random distribution along this phenotypic ranking. One simple version of this strategy selects the top and bottom tails of the distribution, and then tests whether genotypes from these two samples are drawn from a single population. This framework does not require the recovery of phased haplotypes and allows contrasts between large numbers of sites at once. We use a method based on this approach to identify associations between plasma triglyceride level, a risk factor for cardiovascular disease, and multi-site genotypes located in the APOA1/C3/A4/A5 cluster of apolipoprotein genes in unrelated individuals (1,071 African-American females, 780 African-American males, 1,036 European-American females, and 930 European-American males) sampled from four US cities as part of the Coronary Artery Risk Development in Young Adults (CARDIA) study. Method performance is investigated using simulations that model genealogical variation and different genetic architectures. Results indicate that this multi-site test can identify genotype-phenotype associations with reasonable power, including those generated by some simple epistatic models. Genet. Epidemiol . 2006. © 2006 Wiley-Liss, Inc.Peer Reviewedhttp://deepblue.lib.umich.edu/bitstream/2027.42/55790/1/20163_ftp.pd

    A hierarchical Bayesian model for inference of copy number variants and their association to gene expression

    Get PDF
    A number of statistical models have been successfully developed for the analysis of high-throughput data from a single source, but few methods are available for integrating data from different sources. Here we focus on integrating gene expression levels with comparative genomic hybridization (CGH) array measurements collected on the same subjects. We specify a measurement error model that relates the gene expression levels to latent copy number states which, in turn, are related to the observed surrogate CGH measurements via a hidden Markov model. We employ selection priors that exploit the dependencies across adjacent copy number states and investigate MCMC stochastic search techniques for posterior inference. Our approach results in a unified modeling framework for simultaneously inferring copy number variants (CNV) and identifying their significant associations with mRNA transcripts abundance. We show performance on simulated data and illustrate an application to data from a genomic study on human cancer cell lines.Comment: Published in at http://dx.doi.org/10.1214/13-AOAS705 the Annals of Applied Statistics (http://www.imstat.org/aoas/) by the Institute of Mathematical Statistics (http://www.imstat.org
    • …
    corecore