1,446 research outputs found

    Comprehensive analysis of the chromatin landscape in Drosophila melanogaster.

    Get PDF
    Chromatin is composed of DNA and a variety of modified histones and non-histone proteins, which have an impact on cell differentiation, gene regulation and other key cellular processes. Here we present a genome-wide chromatin landscape for Drosophila melanogaster based on eighteen histone modifications, summarized by nine prevalent combinatorial patterns. Integrative analysis with other data (non-histone chromatin proteins, DNase I hypersensitivity, GRO-Seq reads produced by engaged polymerase, short/long RNA products) reveals discrete characteristics of chromosomes, genes, regulatory elements and other functional domains. We find that active genes display distinct chromatin signatures that are correlated with disparate gene lengths, exon patterns, regulatory functions and genomic contexts. We also demonstrate a diversity of signatures among Polycomb targets that include a subset with paused polymerase. This systematic profiling and integrative analysis of chromatin signatures provides insights into how genomic elements are regulated, and will serve as a resource for future experimental investigations of genome structure and function

    Update of the Anopheles gambiae PEST genome assembly

    Get PDF
    BACKGROUND: The genome of Anopheles gambiae, the major vector of malaria, was sequenced and assembled in 2002. This initial genome assembly and analysis made available to the scientific community was complicated by the presence of assembly issues, such as scaffolds with no chromosomal location, no sequence data for the Y chromosome, haplotype polymorphisms resulting in two different genome assemblies in limited regions and contaminating bacterial DNA. RESULTS: Polytene chromosome in situ hybridization with cDNA clones was used to place 15 unmapped scaffolds (sizes totaling 5.34 Mbp) in the pericentromeric regions of the chromosomes and oriented a further 9 scaffolds. Additional analysis by in situ hybridization of bacterial artificial chromosome (BAC) clones placed 1.32 Mbp (5 scaffolds) in the physical gaps between scaffolds on euchromatic parts of the chromosomes. The Y chromosome sequence information (0.18 Mbp) remains highly incomplete and fragmented among 55 short scaffolds. Analysis of BAC end sequences showed that 22 inter-scaffold gaps were spanned by BAC clones. Unmapped scaffolds were also aligned to the chromosome assemblies in silico, identifying regions totaling 8.18 Mbp (144 scaffolds) that are probably represented in the genome project by two alternative assemblies. An additional 3.53 Mbp of alternative assembly was identified within mapped scaffolds. Scaffolds comprising 1.97 Mbp (679 small scaffolds) were identified as probably derived from contaminating bacterial DNA. In total, about 33% of previously unmapped sequences were placed on the chromosomes. CONCLUSION: This study has used new approaches to improve the physical map and assembly of the A. gambiae genome

    Combination of short-read, long-read, and optical mapping assemblies reveals large-scale tandem repeat arrays with population genetic implications

    Get PDF
    Accurate and contiguous genome assembly is key to a comprehensive understanding of the processes shaping genomic diversity and evolution. Yet, it is frequently constrained by constitutive heterochromatin, usually characterized by highly repetitive DNA. As a key feature of genome architecture associated with centromeric and subtelomeric regions, it locally influences meiotic recombination. In this study, we assess the impact of large tandem repeat arrays on the recombination rate landscape in an avian speciation model, the Eurasian crow. We assembled two high-quality genome references using single-molecule real-time sequencing (long-read assembly [LR]) and single-molecule optical maps (optical map assembly [OM]). A three-way comparison including the published short-read assembly (SR) constructed for the same individual allowed assessing assembly properties and pinpointing misassemblies. By combining information from all three assemblies, we characterized 36 previously unidentified large repetitive regions in the proximity of sequence assembly breakpoints, the majority of which contained complex arrays of a 14-kb satellite repeat or its 1.2-kb subunit. Using whole-genome population resequencing data, we estimated the population-scaled recombination rate (ρ) and found it to be significantly reduced in these regions. These findings are consistent with an effect of low recombination in regions adjacent to centromeric or subtelomeric heterochromatin and add to our understanding of the processes generating widespread heterogeneity in genetic diversity and differentiation along the genome. By combining three different technologies, our results highlight the importance of adding a layer of information on genome structure that is inaccessible to each approach independently

    Finishing the euchromatic sequence of the human genome

    Get PDF
    The sequence of the human genome encodes the genetic instructions for human physiology, as well as rich information about human evolution. In 2001, the International Human Genome Sequencing Consortium reported a draft sequence of the euchromatic portion of the human genome. Since then, the international collaboration has worked to convert this draft into a genome sequence with high accuracy and nearly complete coverage. Here, we report the result of this finishing process. The current genome sequence (Build 35) contains 2.85 billion nucleotides interrupted by only 341 gaps. It covers ∼99% of the euchromatic genome and is accurate to an error rate of ∼1 event per 100,000 bases. Many of the remaining euchromatic gaps are associated with segmental duplications and will require focused work with new methods. The near-complete sequence, the first for a vertebrate, greatly improves the precision of biological analyses of the human genome including studies of gene number, birth and death. Notably, the human enome seems to encode only 20,000-25,000 protein-coding genes. The genome sequence reported here should serve as a firm foundation for biomedical research in the decades ahead

    Translational genomics from model species Medicago truncatula to crop legume Trifolium pratense

    Get PDF
    The legume Trifolium pratense (red clover) is an important fodder crop and produces important secondary metabolites. This makes red clover an interesting species. In this thesis, the red clover genome is compared to the legume model species Medicago truncatula, of which the genome sequence is presented. We describe the red clover genome structure and compare it to the Medicago sequence. Thus is shown that although red clover and Medicago are closely related species, their genomes have diverged widely. Further analysis shows that much of the divergence is unique to red clover, not occurring in other clover species. By zooming in on a single rearrangement, a transposable element is found that occurs within the breakpoint region and is widely distributed in the red clover genome, but not in related clover species. Therefore we predict that this transposable element has been involved in the red clover genome rearrangement.­­</p

    Molecular cytogenetic mapping of Cucumis sativus and C. melo using highly repetitive DNA sequences

    Get PDF
    Chromosomes often serve as one of the most important molecular aspects of studying the evolution of species. Indeed, most of the crucial mutations that led to differentiation of species during the evolution have occurred at the chromosomal level. Furthermore, the analysis of pachytene chromosomes appears to be an invaluable tool for the study of evolution due to its effectiveness in chromosome identification and precise physical gene mapping. By applying fluorescence in situ hybridization of 45S rDNA and CsCent1 probes to cucumber pachytene chromosomes, here, we demonstrate that cucumber chromosomes 1 and 2 may have evolved from fusions of ancestral karyotype with chromosome number n= 12. This conclusion is further supported by the centromeric sequence similarity between cucumber and melon, which suggests that these sequences evolved from a common ancestor. It may be after or during speciation that these sequences were specifically amplified, after which they diverged and specific sequence variants were homogenized. Additionally, a structural change on the centromeric region of cucumber chromosome 4 was revealed by fiber-FISH using the mitochondrial-related repetitive sequences, BAC-E38 and CsCent1. These showed the former sequences being integrated into the latter in multiple regions. The data presented here are useful resources for comparative genomics and cytogenetics of Cucumis and, in particular, the ongoing genome sequencing project of cucumbe

    Genome sequence and analysis of the tuber crop potato

    Get PDF
    Potato (Solanum tuberosum L.) is the world’s most important non-grain food crop and is central to global food security. It is clonally propagated, highly heterozygous, autotetraploid, and suffers acute inbreeding depression. Here we use a homozygous doubled-monoploid potato clone to sequence and assemble 86% of the 844-megabase genome. We predict 39,031 protein-coding genes and present evidence for at least two genome duplication events indicative of a palaeopolyploid origin. As the first genome sequence of an asterid, the potato genome reveals 2,642 genes specific to this large angiosperm clade. We also sequenced a heterozygous diploid clone and show that gene presence/absence variants and other potentially deleterious mutations occur frequently and are a likely cause of inbreeding depression. Gene family expansion, tissue-specific expression and recruitment of genes to new pathways contributed to the evolution of tuber development. The potato genome sequence provides a platform for genetic improvement of this vital cro
    corecore