43,013 research outputs found

    A flexible integrative approach based on random forest improves prediction of transcription factor binding sites

    Get PDF
    Transcription factor binding sites (TFBSs) are DNA sequences of 6-15 base pairs. Interaction of these TFBSs with transcription factors (TFs) is largely responsible for most spatiotemporal gene expression patterns. Here, we evaluate to what extent sequence-based prediction of TFBSs can be improved by taking into account the positional dependencies of nucleotides (NPDs) and the nucleotide sequence-dependent structure of DNA. We make use of the random forest algorithm to flexibly exploit both types of information. Results in this study show that both the structural method and the NPD method can be valuable for the prediction of TFBSs. Moreover, their predictive values seem to be complementary, even to the widely used position weight matrix (PWM) method. This led us to combine all three methods. Results obtained for five eukaryotic TFs with different DNA-binding domains show that our method improves classification accuracy for all five eukaryotic TFs compared with other approaches. Additionally, we contrast the results of seven smaller prokaryotic sets with high-quality data and show that with the use of high-quality data we can significantly improve prediction performance. Models developed in this study can be of great use for gaining insight into the mechanisms of TF binding

    Oyster RNA-seq data support the development of Malacoherpesviridae genomics

    Get PDF
    The family of double-stranded DNA (dsDNA) Malacoherpesviridae includes viruses able to infect marine mollusks and detrimental for worldwide aquaculture production. Due to fast-occurring mortality and a lack of permissive cell lines, the available data on the few known Malacoherpesviridae provide only partial support for the study of molecular virus features, life cycle, and evolutionary history. Following thorough data mining of bivalve and gastropod RNA-seq experiments, we used more than five million Malacoherpesviridae reads to improve the annotation of viral genomes and to characterize viral InDels, nucleotide stretches, and SNPs. Both genome and protein domain analyses confirmed the evolutionary diversification and gene uniqueness of known Malacoherpesviridae. However, the presence of Malacoherpesviridae-like sequences integrated within genomes of phylogenetically distant invertebrates indicates broad diffusion of these viruses and indicates the need for confirmatory investigations. The manifest co-occurrence of OsHV-1 genotype variants in single RNA-seq samples of Crassostrea gigas provide further support for the Malacoherpesviridae diversification. In addition to simple sequence motifs inter-punctuating viral ORFs, recombination-inducing sequences were found to be enriched in the OsHV-1 and AbHV1-AUS genomes. Finally, the highly correlated expression of most viral ORFs in multiple oyster samples is consistent with the burst of viral proteins during the lytic phase

    Intragenic homogenization and multiple copies of prey-wrapping silk genes in Argiope garden spiders.

    Get PDF
    BackgroundSpider silks are spectacular examples of phenotypic diversity arising from adaptive molecular evolution. An individual spider can produce an array of specialized silks, with the majority of constituent silk proteins encoded by members of the spidroin gene family. Spidroins are dominated by tandem repeats flanked by short, non-repetitive N- and C-terminal coding regions. The remarkable mechanical properties of spider silks have been largely attributed to the repeat sequences. However, the molecular evolutionary processes acting on spidroin terminal and repetitive regions remain unclear due to a paucity of complete gene sequences and sampling of genetic variation among individuals. To better understand spider silk evolution, we characterize a complete aciniform spidroin gene from an Argiope orb-weaving spider and survey aciniform gene fragments from congeneric individuals.ResultsWe present the complete aciniform spidroin (AcSp1) gene from the silver garden spider Argiope argentata (Aar_AcSp1), and document multiple AcSp1 loci in individual genomes of A. argentata and the congeneric A. trifasciata and A. aurantia. We find that Aar_AcSp1 repeats have >98% pairwise nucleotide identity. By comparing AcSp1 repeat amino acid sequences between Argiope species and with other genera, we identify regions of conservation over vast amounts of evolutionary time. Through a PCR survey of individual A. argentata, A. trifasciata, and A. aurantia genomes, we ascertain that AcSp1 repeats show limited variation between species whereas terminal regions are more divergent. We also find that average dN/dS across codons in the N-terminal, repetitive, and C-terminal encoding regions indicate purifying selection that is strongest in the N-terminal region.ConclusionsUsing the complete A. argentata AcSp1 gene and spidroin genetic variation between individuals, this study clarifies some of the molecular evolutionary processes underlying the spectacular mechanical attributes of aciniform silk. It is likely that intragenic concerted evolution and functional constraints on A. argentata AcSp1 repeats result in extreme repeat homogeneity. The maintenance of multiple AcSp1 encoding loci in Argiope genomes supports the hypothesis that Argiope spiders require rapid and efficient protein production to support their prolific use of aciniform silk for prey-wrapping and web-decorating. In addition, multiple gene copies may represent the early stages of spidroin diversification

    A fast and cost-effective approach to develop and map EST-SSR markers: oak as a case study

    Get PDF
    Background: Expressed Sequence Tags (ESTs) are a source of simple sequence repeats (SSRs) that can be used to develop molecular markers for genetic studies. The availability of ESTs for Quercus robur and Quercus petraea provided a unique opportunity to develop microsatellite markers to accelerate research aimed at studying adaptation of these long-lived species to their environment. As a first step toward the construction of a SSR-based linkage map of oak for quantitative trait locus (QTL) mapping, we describe the mining and survey of EST-SSRs as well as a fast and cost-effective approach (bin mapping) to assign these markers to an approximate map position. We also compared the level of polymorphism between genomic and EST-derived SSRs and address the transferability of EST-SSRs in Castanea sativa (chestnut). Results: A catalogue of 103,000 Sanger ESTs was assembled into 28,024 unigenes from which 18.6% presented one or more SSR motifs. More than 42% of these SSRs corresponded to trinucleotides. Primer pairs were designed for 748 putative unigenes. Overall 37.7% (283) were found to amplify a single polymorphic locus in a reference fullsib pedigree of Quercus robur. The usefulness of these loci for establishing a genetic map was assessed using a bin mapping approach. Bin maps were constructed for the male and female parental tree for which framework linkage maps based on AFLP markers were available. The bin set consisting of 14 highly informative offspring selected based on the number and position of crossover sites. The female and male maps comprised 44 and 37 bins, with an average bin length of 16.5 cM and 20.99 cM, respectively. A total of 256 EST-SSRs were assigned to bins and their map position was further validated by linkage mapping. EST-SSRs were found to be less polymorphic than genomic SSRs, but their transferability rate to chestnut, a phylogenetically related species to oak, was higher. Conclusion: We have generated a bin map for oak comprising 256 EST-SSRs. This resource constitutes a first step toward the establishment of a gene-based map for this genus that will facilitate the dissection of QTLs affecting complex traits of ecological importance

    Fundamental principles in drawing inference from sequence analysis

    No full text
    Individual life courses are dynamic and can be represented as a sequence of states for some portion of their experiences. More generally, study of such sequences has been made in many fields around social science; for example, sociology, linguistics, psychology, and the conceptualisation of subjects progressing through a sequence of states is common. However, many models and sets of data allow only for the treatment of aggregates or transitions, rather than interpreting whole sequences. The temporal aspect of the analysis is fundamental to any inference about the evolution of the subjects but assumptions about time are not normally made explicit. Moreover, without a clear idea of what sequences look like, it is impossible to determine when something is not seen whether it was not actually there. Some principles are proposed which link the ideas of sequences, hypothesis, analytical framework, categorisation and representation; each one being underpinned by the consideration of time. To make inferences about sequences, one needs to: understand what these sequences represent; the hypothesis and assumptions that can be derived about sequences; identify the categories within the sequences; and data representation at each stage. These ideas are obvious in themselves but they are interlinked, imposing restrictions on each other and on the inferences which can be draw
    corecore