3,306 research outputs found

    Computational Methods for Comparative Non-coding RNA Analysis: from Secondary Structures to Tertiary Structures

    Get PDF
    Unlike message RNAs (mRNAs) whose information is encoded in the primary sequences, the cellular roles of non-coding RNAs (ncRNAs) originate from the structures. Therefore studying the structural conservation in ncRNAs is important to yield an in-depth understanding of their functionalities. In the past years, many computational methods have been proposed to analyze the common structural patterns in ncRNAs using comparative methods. However, the RNA structural comparison is not a trivial task, and the existing approaches still have numerous issues in efficiency and accuracy. In this dissertation, we will introduce a suite of novel computational tools that extend the classic models for ncRNA secondary and tertiary structure comparisons. For RNA secondary structure analysis, we first developed a computational tool, named PhyloRNAalifold, to integrate the phylogenetic information into the consensus structural folding. The underlying idea of this algorithm is that the importance of a co-varying mutation should be determined by its position on the phylogenetic tree. By assigning high scores to the critical covariances, the prediction of RNA secondary structure can be more accurate. Besides structure prediction, we also developed a computational tool, named ProbeAlign, to improve the efficiency of genome-wide ncRNA screening by using high-throughput RNA structural probing data. It treats the chemical reactivities embedded in the probing information as pairing attributes of the searching targets. This approach can avoid the time-consuming base pair matching in the secondary structure alignment. The application of ProbeAlign to the FragSeq datasets shows its capability of genome-wide ncRNAs analysis. For RNA tertiary structure analysis, we first developed a computational tool, named STAR3D, to find the global conservation in RNA 3D structures. STAR3D aims at finding the consensus of stacks by using 2D topology and 3D geometry together. Then, the loop regions can be ordered and aligned according to their relative positions in the consensus. This stack-guided alignment method adopts the divide-and-conquer strategy into RNA 3D structural alignment, which has improved its efficiency dramatically. Furthermore, we also have clustered all loop regions in non-redundant RNA 3D structures to de novo detect plausible RNA structural motifs. The computational pipeline, named RNAMSC, was extended to handle large-scale PDB datasets, and solid downstream analysis was performed to ensure the clustering results are valid and easily to be applied to further research. The final results contain many interesting variations of known motifs, such as GNAA tetraloop, kink-turn, sarcin-ricin and t-loops. We also discovered novel functional motifs that conserved in a wide range of ncRNAs, including ribosomal RNA, sgRNA, SRP RNA, GlmS riboswitch and twister ribozyme

    Development of ListeriaBase and comparative analysis of Listeria monocytogenes

    Get PDF
    Background: Listeria consists of both pathogenic and non-pathogenic species. Reports of similarities between the genomic content between some pathogenic and non-pathogenic species necessitates the investigation of these species at the genomic level to understand the evolution of virulence-associated genes. With Listeria genome data growing exponentially, comparative genomic analysis may give better insights into evolution, genetics and phylogeny of Listeria spp., leading to better management of the diseases caused by them. Description: With this motivation, we have developed ListeriaBase, a web Listeria genomic resource and analysis platform to facilitate comparative analysis of Listeria spp. ListeriaBase currently houses 850,402 protein-coding genes, 18,113 RNAs and 15,576 tRNAs from 285 genome sequences of different Listeria strains. An AJAX-based real time search system implemented in ListeriaBase facilitates searching of this huge genomic data. Our in-house designed comparative analysis tools such as Pairwise Genome Comparison (PGC) tool allowing comparison between two genomes, Pathogenomics Profiling Tool (PathoProT) for comparing the virulence genes, and ListeriaTree for phylogenic classification, were customized and incorporated in ListeriaBase facilitating comparative genomic analysis of Listeria spp. Interestingly, we identified a unique genomic feature in the L. monocytogenes genomes in our analysis. The Auto protein sequences of the serotype 4 and the non-serotype 4 strains of L. monocytogenes possessed unique sequence signatures that can differentiate the two groups. We propose that the aut gene may be a potential gene marker for differentiating the serotype 4 strains from other serotypes of L. monocytogenes. Conclusions: ListeriaBase is a useful resource and analysis platform that can facilitate comparative analysis of Listeria for the scientific communities. We have successfully demonstrated some key utilities of ListeriaBase. The knowledge that we obtained in the analyses of L. monocytogenes may be important for functional works of this human pathogen in future. ListeriaBase is currently available at http://listeria.um.edu.my

    The Qphyl System: a web-based interactive system for phylogenetic analysis

    Get PDF
    Phylogenetic tree reconstruction is a prominent problem in computational biology. Currently, all computational methods have their limitations and work well only for simple problems of small size. No existing method can guarantee that trees constructed for real-world problems are true phylogenetic trees for large and complex problems mainly because the existing computational models are not very biologically realistic. It has become a serious issue for many important real-life applications which often desire accurate results from phylogenetic analysis. Thus, it is very crucial to effectively incorporate multi-disciplinary analyses and synthesize results from various sources when answering real-life questions. In this thesis, a novel web-based phylogeny reconstruction system with a real-time interactive environment, called Qphyl (short for quartet-based phylogenetic analysis) is introduced. The Qphyl system uses a new interactive approach to enable biologists to greatly improve the final results through effectively dynamic interaction with the computation, e.g., to move the computation back and forth to different stages so users can check the intermediate results, compare results from different methods and carry out certain manual refinements using their biological domain-specific knowledge in the decision making on how a tree should be reconstructed. Currently the alpha version of this web-based interactive system has been released and accessible through the URL: http://ww-test.it.usyd.edu.au/sogrid/qphyl/

    tRNA functional signatures classify plastids as late-branching cyanobacteria.

    Get PDF
    BackgroundEukaryotes acquired the trait of oxygenic photosynthesis through endosymbiosis of the cyanobacterial progenitor of plastid organelles. Despite recent advances in the phylogenomics of Cyanobacteria, the phylogenetic root of plastids remains controversial. Although a single origin of plastids by endosymbiosis is broadly supported, recent phylogenomic studies are contradictory on whether plastids branch early or late within Cyanobacteria. One underlying cause may be poor fit of evolutionary models to complex phylogenomic data.ResultsUsing Posterior Predictive Analysis, we show that recently applied evolutionary models poorly fit three phylogenomic datasets curated from cyanobacteria and plastid genomes because of heterogeneities in both substitution processes across sites and of compositions across lineages. To circumvent these sources of bias, we developed CYANO-MLP, a machine learning algorithm that consistently and accurately phylogenetically classifies ("phyloclassifies") cyanobacterial genomes to their clade of origin based on bioinformatically predicted function-informative features in tRNA gene complements. Classification of cyanobacterial genomes with CYANO-MLP is accurate and robust to deletion of clades, unbalanced sampling, and compositional heterogeneity in input tRNA data. CYANO-MLP consistently classifies plastid genomes into a late-branching cyanobacterial sub-clade containing single-cell, starch-producing, nitrogen-fixing ecotypes, consistent with metabolic and gene transfer data.ConclusionsPhylogenomic data of cyanobacteria and plastids exhibit both site-process heterogeneities and compositional heterogeneities across lineages. These aspects of the data require careful modeling to avoid bias in phylogenomic estimation. Furthermore, we show that amino acid recoding strategies may be insufficient to mitigate bias from compositional heterogeneities. However, the combination of our novel tRNA-specific strategy with machine learning in CYANO-MLP appears robust to these sources of bias with high accuracy in phyloclassification of cyanobacterial genomes. CYANO-MLP consistently classifies plastids as late-branching Cyanobacteria, consistent with independent evidence from signature-based approaches and some previous phylogenetic studies

    High Performance Implementation of Planted Motif Problem using Suffix trees

    Get PDF
    In this paper we present a high performance implementation of suffix tree based solution to the planted motif problem on two different parallel architectures: NVIDIA GPU and Intel Multicore machines. An (l,d) planted motif problem(PMP) is defined as: Given a sequence of n DNA sequences, each of length L, find M, the set of sequences(or motifs) of length l which have atleast one d-neighbor in each of the n sequences. Here, a d-neighbor of a sequence is a sequence of same length that differs in at-most d positions. PMP is a well studied problem in computational biology. It is useful in developing methods for finding transcription factor binding sites, sequence classification and for building phylogenetic trees. The problem is computationally challenging to solve, for example a (19,7) PMP takes 9.9 hours on a sequential machine. Many approaches to solve planted motif problem can be found in literature. One approach is based on use of suffix tree data structure. Though suffix tree based methods are the most efficient ones for solving large planted motif problems on sequential machines, they are quite difficult to parallelize. We present suffix tree based parallel solutions for PMP on NVIDIA GPU and Intel Multicore architectures that are efficient and scalable. The solutions are based on a suffix tree algorithm previously presented but use extensive adaptation to individual architectures to ensure that the implementations work efficiently and scale well

    Biogenesis of the inner membrane complex is dependent on vesicular transport by the alveolate specific GTPase Rab11B

    Get PDF
    Apicomplexan parasites belong to a recently recognised group of protozoa referred to as Alveolata. These protists contain membranous sacs (alveoli) beneath the plasma membrane, termed the Inner Membrane Complex (IMC) in the case of Apicomplexa. During parasite replication the IMC is formed de novo within the mother cell in a process described as internal budding. We hypothesized that an alveolate specific factor is involved in the specific transport of vesicles from the Golgi to the IMC and identified the small GTPase Rab11B as an alveolate specific Rab-GTPase that localises to the growing end of the IMC during replication of Toxoplasma gondii. Conditional interference with Rab11B function leads to a profound defect in IMC biogenesis, indicating that Rab11B is required for the transport of Golgi derived vesicles to the nascent IMC of the daughter cell. Curiously, a block in IMC biogenesis did not affect formation of sub-pellicular microtubules, indicating that IMC biogenesis and formation of sub-pellicular microtubules is not mechanistically linked. We propose a model where Rab11B specifically transports vesicles derived from the Golgi to the immature IMC of the growing daughter parasites

    A new multi locus variable number of tandem repeat analysis scheme for epidemiological surveillance of Xanthomonas vasicola pv. musacearum, the plant pathogen causing bacterial wilt on banana and enset

    Get PDF
    Xanthomonas vasicola pv. musacearum (Xvm) which causes Xanthomonas wilt (XW) on banana (Musa accuminata x balbisiana) and enset (Ensete ventricosum), is closely related to the species Xanthomonas vasicola that contains the pathovars vasculorum (Xvv) and holcicola (Xvh), respectively pathogenic to sugarcane and sorghum. Xvm is considered a monomorphic bacterium whose intra-pathovar diversity remains poorly understood. With the sudden emergence of Xvm within east and central Africa coupled with the unknown origin of one of the two sublineages suggested for Xvm, attention has shifted to adapting technologies that focus on identifying the origin and distribution of the genetic diversity within this pathogen. Although microbiological and conventional molecular diagnostics have been useful in pathogen identification. Recent advances have ushered in an era of genomic epidemiology that aids in characterizing monomorphic pathogens. To unravel the origin and pathways of the recent emergence of XW in Eastern and Central Africa, there was a need for a genotyping tool adapted for molecular epidemiology. Multi-Locus Variable Number of Tandem Repeat Analysis (MLVA) is able to resolve the evolutionary patterns and invasion routes of a pathogen. In this study, we identified microsatellite loci from nine published Xvm genome sequences. Of the 36 detected microsatellite loci, 21 were selected for primer design and 19 determined to be highly typeable, specific, reproducible and polymorphic with two- to four- alleles per locus on a sub-collection. The 19 markers were multiplexed and applied to genotype 335 Xvm strains isolated from seven countries over several years. The microsatellite markers grouped the Xvm collection into three clusters; with two similar to the SNP-based sublineages 1 and 2 and a new cluster 3, revealing an unknown diversity in Ethiopia. Five of the 19 markers had alleles present in both Xvm and Xanthomonas vasicola pathovars holcicola and vasculorum, supporting the phylogenetic closeliness of these three pathovars. Thank to the public availability of the haplotypes on the MLVABank database, this highly reliable and polymorphic genotyping tool can be further used in a transnational surveillance network to monitor the spread and evolution of XW throughout Africa.. It will inform and guide management of Xvm both in banana-based and enset-based cropping systems. Due to the suitability of MLVA-19 markers for population genetic analyses, this genotyping tool will also be used in future microevolution studies
    • …
    corecore