706 research outputs found

    Statistical Viewer: a tool to upload and integrate linkage and association data as plots displayed within the Ensembl genome browser

    Get PDF
    BACKGROUND: To facilitate efficient selection and the prioritization of candidate complex disease susceptibility genes for association analysis, increasingly comprehensive annotation tools are essential to integrate, visualize and analyze vast quantities of disparate data generated by genomic screens, public human genome sequence annotation and ancillary biological databases. We have developed a plug-in package for Ensembl called "Statistical Viewer" that facilitates the analysis of genomic features and annotation in the regions of interest defined by linkage analysis. RESULTS: Statistical Viewer is an add-on package to the open-source Ensembl Genome Browser and Annotation System that displays disease study-specific linkage and/or association data as 2 dimensional plots in new panels in the context of Ensembl's Contig View and Cyto View pages. An enhanced upload server facilitates the upload of statistical data, as well as additional feature annotation to be displayed in DAS tracts, in the form of Excel Files. The Statistical View panel, drawn directly under the ideogram, illustrates lod score values for markers from a study of interest that are plotted against their position in base pairs. A module called "Get Map" easily converts the genetic locations of markers to genomic coordinates. The graph is placed under the corresponding ideogram features a synchronized vertical sliding selection box that is seamlessly integrated into Ensembl's Contig- and Cyto- View pages to choose the region to be displayed in Ensembl's "Overview" and "Detailed View" panels. To resolve Association and Fine mapping data plots, a "Detailed Statistic View" plot corresponding to the "Detailed View" may be displayed underneath. CONCLUSION: Features mapping to regions of linkage are accentuated when Statistic View is used in conjunction with the Distributed Annotation System (DAS) to display supplemental laboratory information such as differentially expressed disease genes in private data tracks. Statistic View is a novel and powerful visual feature that enhances Ensembl's utility as valuable resource for integrative genomic-based approaches to the identification of candidate disease susceptibility genes. At present there are no other tools that provide for the visualization of 2-dimensional plots of quantitative data scores against genomic coordinates in the context of a primary public genome annotation browser

    A noise-reduction GWAS analysis implicates altered regulation of neurite outgrowth and guidance in autism

    Get PDF
    <p>Abstract</p> <p>Background</p> <p>Genome-wide Association Studies (GWAS) have proved invaluable for the identification of disease susceptibility genes. However, the prioritization of candidate genes and regions for follow-up studies often proves difficult due to false-positive associations caused by statistical noise and multiple-testing. In order to address this issue, we propose the novel GWAS noise reduction (GWAS-NR) method as a way to increase the power to detect true associations in GWAS, particularly in complex diseases such as autism.</p> <p>Methods</p> <p>GWAS-NR utilizes a linear filter to identify genomic regions demonstrating correlation among association signals in multiple datasets. We used computer simulations to assess the ability of GWAS-NR to detect association against the commonly used joint analysis and Fisher's methods. Furthermore, we applied GWAS-NR to a family-based autism GWAS of 597 families and a second existing autism GWAS of 696 families from the Autism Genetic Resource Exchange (AGRE) to arrive at a compendium of autism candidate genes. These genes were manually annotated and classified by a literature review and functional grouping in order to reveal biological pathways which might contribute to autism aetiology.</p> <p>Results</p> <p>Computer simulations indicate that GWAS-NR achieves a significantly higher classification rate for true positive association signals than either the joint analysis or Fisher's methods and that it can also achieve this when there is imperfect marker overlap across datasets or when the closest disease-related polymorphism is not directly typed. In two autism datasets, GWAS-NR analysis resulted in 1535 significant linkage disequilibrium (LD) blocks overlapping 431 unique reference sequencing (RefSeq) genes. Moreover, we identified the nearest RefSeq gene to the non-gene overlapping LD blocks, producing a final candidate set of 860 genes. Functional categorization of these implicated genes indicates that a significant proportion of them cooperate in a coherent pathway that regulates the directional protrusion of axons and dendrites to their appropriate synaptic targets.</p> <p>Conclusions</p> <p>As statistical noise is likely to particularly affect studies of complex disorders, where genetic heterogeneity or interaction between genes may confound the ability to detect association, GWAS-NR offers a powerful method for prioritizing regions for follow-up studies. Applying this method to autism datasets, GWAS-NR analysis indicates that a large subset of genes involved in the outgrowth and guidance of axons and dendrites is implicated in the aetiology of autism.</p

    Combinatorial Mismatch Scan (CMS) for loci associated with dementia in the Amish

    Get PDF
    BACKGROUND: Population heterogeneity may be a significant confounding factor hampering detection and verification of late onset Alzheimer's disease (LOAD) susceptibility genes. The Amish communities located in Indiana and Ohio are relatively isolated populations that may have increased power to detect disease susceptibility genes. METHODS: We recently performed a genome scan of dementia in this population that detected several potential loci. However, analyses of these data are complicated by the highly consanguineous nature of these Amish pedigrees. Therefore we applied the Combinatorial Mismatch Scanning (CMS) method that compares identity by state (IBS) (under the presumption of identity by descent (IBD)) sharing in distantly related individuals from such populations where standard linkage and association analyses are difficult to implement. CMS compares allele sharing between individuals in affected and unaffected groups from founder populations. Comparisons between cases and controls were done using two Fisher's exact tests, one testing for excess in IBS allele frequency and the other testing for excess in IBS genotype frequency for 407 microsatellite markers. RESULTS: In all, 13 dementia cases and 14 normal controls were identified who were not related at least through the grandparental generation. The examination of allele frequencies identified 24 markers (6%) nominally (p ≤ 0.05) associated with dementia; the most interesting (empiric p ≤ 0.005) markers were D3S1262, D5S211, and D19S1165. The examination of genotype frequencies identified 21 markers (5%) nominally (p ≤ 0.05) associated with dementia; the most significant markers were both located on chromosome 5 (D5S1480 and D5S211). Notably, one of these markers (D5S211) demonstrated differences (empiric p ≤ 0.005) under both tests. CONCLUSION: Our results provide the initial groundwork for identifying genes involved in late-onset Alzheimer's disease within the Amish community. Genes identified within this isolated population will likely play a role in a subset of late-onset AD cases across more general populations. Regions highlighted by markers demonstrating suggestive allelic and/or genotypic differences will be the focus of more detailed examination to characterize their involvement in dementia

    Linkage analyses in Caribbean Hispanic families identify novel loci associated with familial late-onset Alzheimer's disease

    Get PDF
    INTRODUCTION: We performed linkage analyses in Caribbean Hispanic families with multiple late-onset Alzheimer's disease (LOAD) cases to identify regions that may contain disease causative variants. METHODS: We selected 67 LOAD families to perform genome-wide linkage scan. Analysis of the linked regions was repeated using the entire sample of 282 families. Validated chromosomal regions were analyzed using joint linkage and association. RESULTS: We identified 26 regions linked to LOAD (HLOD ≥3.6). We validated 13 of the regions (HLOD ≥2.5) using the entire family sample. The strongest signal was at 11q12.3 (rs2232932: HLODmax = 4.7, Pjoint = 6.6 × 10(-6)), a locus located ∼2 Mb upstream of the membrane-spanning 4A gene cluster. We additionally identified a locus at 7p14.3 (rs10255835: HLODmax = 4.9, Pjoint = 1.2 × 10(-5)), a region harboring genes associated with the nervous system (GARS, GHRHR, and NEUROD6). DISCUSSION: Future sequencing efforts should focus on these regions because they may harbor familial LOAD causative mutations

    Copy Number Variants in Extended Autism Spectrum Disorder Families Reveal Candidates Potentially Involved in Autism Risk

    Get PDF
    Copy number variations (CNVs) are a major cause of genetic disruption in the human genome with far more nucleotides being altered by duplications and deletions than by single nucleotide polymorphisms (SNPs). In the multifaceted etiology of autism spectrum disorders (ASDs), CNVs appear to contribute significantly to our understanding of the pathogenesis of this complex disease. A unique resource of 42 extended ASD families was genotyped for over 1 million SNPs to detect CNVs that may contribute to ASD susceptibility. Each family has at least one avuncular or cousin pair with ASD. Families were then evaluated for co-segregation of CNVs in ASD patients. We identified a total of five deletions and seven duplications in eleven families that co-segregated with ASD. Two of the CNVs overlap with regions on 7p21.3 and 15q24.1 that have been previously reported in ASD individuals and two additional CNVs on 3p26.3 and 12q24.32 occur near regions associated with schizophrenia. These findings provide further evidence for the involvement of ICA1 and NXPH1 on 7p21.3 in ASD susceptibility and highlight novel ASD candidates, including CHL1, FGFBP3 and POUF41. These studies highlight the power of using extended families for gene discovery in traits with a complex etiology

    An X chromosome-wide association study in autism families identifies TBL1X as a novel autism spectrum disorder candidate gene in males

    Get PDF
    <p>Abstract</p> <p>Background</p> <p>Autism spectrum disorder (ASD) is a complex neurodevelopmental disorder with a strong genetic component. The skewed prevalence toward males and evidence suggestive of linkage to the X chromosome in some studies suggest the presence of X-linked susceptibility genes in people with ASD.</p> <p>Methods</p> <p>We analyzed genome-wide association study (GWAS) data on the X chromosome in three independent autism GWAS data sets: two family data sets and one case-control data set. We performed meta- and joint analyses on the combined family and case-control data sets. In addition to the meta- and joint analyses, we performed replication analysis by using the two family data sets as a discovery data set and the case-control data set as a validation data set.</p> <p>Results</p> <p>One SNP, rs17321050, in the transducin β-like 1X-linked (<it>TBL1X</it>) gene [OMIM:300196] showed chromosome-wide significance in the meta-analysis (<it>P </it>value = 4.86 × 10<sup>-6</sup>) and joint analysis (<it>P </it>value = 4.53 × 10<sup>-6</sup>) in males. The SNP was also close to the replication threshold of 0.0025 in the discovery data set (<it>P </it>= 5.89 × 10<sup>-3</sup>) and passed the replication threshold in the validation data set (<it>P </it>= 2.56 × 10<sup>-4</sup>). Two other SNPs in the same gene in linkage disequilibrium with rs17321050 also showed significance close to the chromosome-wide threshold in the meta-analysis.</p> <p>Conclusions</p> <p><it>TBL1X </it>is in the Wnt signaling pathway, which has previously been implicated as having a role in autism. Deletions in the Xp22.2 to Xp22.3 region containing <it>TBL1X </it>and surrounding genes are associated with several genetic syndromes that include intellectual disability and autistic features. Our results, based on meta-analysis, joint analysis and replication analysis, suggest that <it>TBL1X </it>may play a role in ASD risk.</p

    Evidence of novel finescale structural variation at autism spectrum disorder candidate loci

    Get PDF
    Background: Autism spectrum disorders (ASD) represent a group of neurodevelopmental disorders characterized by a core set of social-communicative and behavioral impairments. Gamma-aminobutyric acid (GABA) is the major inhibitory neurotransmitter in the brain, acting primarily via the GABA receptors (GABR). Multiple lines of evidence, including altered GABA and GABA receptor expression in autistic patients, indicate that the GABAergic system may be involved in the etiology of autism. Methods: As copy number variations (CNVs), particularly rare and de novo CNVs, have now been implicated in ASD risk, we examined the GABA receptors and genes in related pathways for structural variation that may be associated with autism. We further extended our candidate gene set to include 19 genes and regions that had either been directly implicated in the autism literature or were directly related (via function or ancestry) to these primary candidates. For the high resolution CNV screen we employed custom-designed 244 k comparative genomic hybridization (CGH) arrays. Collectively, our probes spanned a total of 11 Mb of GABA-related and additional candidate regions with a density of approximately one probe every 200 nucleotides, allowing a theoretical resolution for detection of CNVs of approximately 1 kb or greater on average. One hundred and sixty-eight autism cases and 149 control individuals were screened for structural variants. Prioritized CNV events were confirmed using quantitative PCR, and confirmed loci were evaluated on an additional set of 170 cases and 170 control individuals that were not included in the original discovery set. Loci that remained interesting were subsequently screened via quantitative PCR on an additional set of 755 cases and 1,809 unaffected family members. Results: Results include rare deletions in autistic individuals at JAKMIP1, NRXN1, Neuroligin4Y, OXTR, and ABAT. Common insertion/deletion polymorphisms were detected at several loci, including GABBR2 and NRXN3. Overall, statistically significant enrichment in affected vs. unaffected individuals was observed for NRXN1 deletions. Conclusions: These results provide additional support for the role of rare structural variation in ASD

    Targeted massively parallel sequencing of autism spectrum disorder-associated genes in a case control cohort reveals rare loss-of-function risk variants

    Get PDF
    BACKGROUND: Autism spectrum disorder (ASD) is highly heritable, yet genome-wide association studies (GWAS), copy number variation screens, and candidate gene association studies have found no single factor accounting for a large percentage of genetic risk. ASD trio exome sequencing studies have revealed genes with recurrent de novo loss-of-function variants as strong risk factors, but there are relatively few recurrently affected genes while as many as 1000 genes are predicted to play a role. As such, it is critical to identify the remaining rare and low-frequency variants contributing to ASD. METHODS: We have utilized an approach of prioritization of genes by GWAS and follow-up with massively parallel sequencing in a case-control cohort. Using a previously reported ASD noise reduction GWAS analyses, we prioritized 837 RefSeq genes for custom targeting and sequencing. We sequenced the coding regions of those genes in 2071 ASD cases and 904 controls of European white ancestry. We applied comprehensive annotation to identify single variants which could confer ASD risk and also gene-based association analysis to identify sets of rare variants associated with ASD. RESULTS: We identified a significant over-representation of rare loss-of-function variants in genes previously associated with ASD, including a de novo premature stop variant in the well-established ASD candidate gene RBFOX1. Furthermore, ASD cases were more likely to have two damaging missense variants in candidate genes than controls. Finally, gene-based rare variant association implicates genes functioning in excitatory neurotransmission and neurite outgrowth and guidance pathways including CACNAD2, KCNH7, and NRXN1. CONCLUSIONS: We find suggestive evidence that rare variants in synaptic genes are associated with ASD and that loss-of-function mutations in ASD candidate genes are a major risk factor, and we implicate damaging mutations in glutamate signaling receptors and neuronal adhesion and guidance molecules. Furthermore, the role of de novo mutations in ASD remains to be fully investigated as we identified the first reported protein-truncating variant in RBFOX1 in ASD. Overall, this work, combined with others in the field, suggests a convergence of genes and molecular pathways underlying ASD etiology. ELECTRONIC SUPPLEMENTARY MATERIAL: The online version of this article (doi:10.1186/s13229-015-0034-z) contains supplementary material, which is available to authorized users

    A comparative analysis of the information content in long and short SAGE libraries

    Get PDF
    BACKGROUND: Serial Analysis of Gene Expression (SAGE) is a powerful tool to determine gene expression profiles. Two types of SAGE libraries, ShortSAGE and LongSAGE, are classified based on the length of the SAGE tag (10 vs. 17 basepairs). LongSAGE libraries are thought to be more useful than ShortSAGE libraries, but their information content has not been widely compared. To dissect the differences between these two types of libraries, we utilized four libraries (two LongSAGE and two ShortSAGE libraries) generated from the hippocampus of Alzheimer and control samples. In addition, we generated two additional short SAGE libraries, the truncated long SAGE libraries (tSAGE), from LongSAGE libraries by deleting seven 5' basepairs from each LongSAGE tag. RESULTS: One problem that occurred in the SAGE study is that individual tags may have matched to multiple different genes – due to the short length of a tag. We found that the LongSAGE tag maps up to 15 UniGene clusters, while the ShortSAGE and tSAGE tags map up to 279 UniGene clusters. Both long and short SAGE libraries exhibit a large number of orphan tags (no gene information in UniGene), implying the limitation of the UniGene database. Among 100 orphan LongSAGE tags, the complete sequences (17 basepairs) of nine orphan tags match to 17 genomic sequences; four of the orphan tags match to a single genomic sequence. Our data show the potential to resolve 4–9% of orphan LongSAGE tags. Finally, among 400 tSAGE tags showing significant differential expression between AD and control, 79 tags (19.8%) were derived from multiple non-significant LongSAGE tags, implying the false positive results. CONCLUSION: Our data show that LongSAGE tags have high specificity in gene mapping compared to ShortSAGE tags. LongSAGE tags show an advantage over ShortSAGE in identifying novel genes by BLAST analysis. Most importantly, the chances of obtaining false positive results are higher for ShortSAGE than LongSAGE libraries due to their specificity in gene mapping. Therefore, it is recommended that the number of corresponding UniGene clusters (gene or ESTs) of a tag for prioritizing the significant results be considered
    corecore