999 research outputs found
Exploring data processing strategies in NGS target enrichment to disentangle radiations in the tribe Cardueae (Compositae)
Target enrichment is a cost-effective sequencing technique that holds promise for elucidating evolutionary relationships in fast-evolving lineages. However, potential biases and impact of bioinformatic sequence treatments in phylogenetic inference have not been thoroughly explored yet. Here, we investigate this issue with an ultimate goal to shed light into a highly diversified group of Compositae (Asteraceae) constituted by four main genera: Arctium, Cousinia, Saussurea, and Jurinea. Specifically, we compared sequence data extraction methods implemented in two easy-to-use workflows, PHYLUCE and HybPiper, and assessed the impact of two filtering practices intended to reduce phylogenetic noise. In addition, we compared two phylogenetic inference methods: (1) the concatenation approach, in which all loci were concatenated in a supermatrix; and (2) the coalescence approach, in which gene trees were produced independently and then used to construct a species tree under coalescence assumptions. Here we confirm the usefulness of the set of 1061 COS targets (a nuclear conserved orthology loci set developed for the Compositae) across a variety of taxonomic levels. Intergeneric relationships were completely resolved: there are two sister groups, Arctium-Cousinia and Saussurea-Jurinea, which are in agreement with a morphological hypothesis. Intrageneric relationships among species of Arctium, Cousinia, and Saussurea are also well defined. Conversely, conflicting species relationships remain for Jurinea. Methodological choices significantly affected phylogenies in terms of topology, branch length, and support. Across all analyses, the phylogeny obtained using HybPiper and the strictest scheme of removing fast-evolving sites was estimated as the optimal. Regarding methodological choices, we conclude that (1) trees obtained under the coalescence approach are topologically more congruent between them than those inferred using the concatenation approach; (2) refining treatments only improved support values under the concatenation approach; and (3) branch support values are maximized when fast-evolving sites are removed in the concatenation approach, and when a higher number of loci is analyzed in the coalescence approach.Peer reviewe
Comparative genomic analysis and molecular examination of the diversity of enterotoxigenic Escherichia coli isolates from Chile
Enterotoxigenic Escherichia coli (ETEC) is one of the most common diarrheal pathogens in the low- and middle-income regions of the world, however a systematic examination of the genomic content of isolates from Chile has not yet been undertaken. Whole genome sequencing and comparative analysis of a collection of 125 ETEC isolates from three geographic locations in Chile, allowed the interrogation of phylogenomic groups, sequence types and genes specific to isolates from the different geographic locations. A total of 80.8% (101/125) of the ETEC isolates were identified in E. coli phylogroup A, 15.2% (19/125) in phylogroup B, and 4.0% (5/125) in phylogroup E. The over-representation of genomes in phylogroup A was significantly different from other global ETEC genomic studies. The Chilean ETEC isolates could be further subdivided into sub-clades similar to previously defined global ETEC reference lineages that had conserved multi-locus sequence types and toxin profiles. Comparison of the gene content of the Chilean ETEC identified genes that were unique based on geographic location within Chile, phylogenomic classifications or sequence type. Completion of a limited number of genomes provided insight into the ETEC plasmid content, which is conserved in some phylogenomic groups and not conserved in others. These findings suggest that the Chilean ETEC isolates contain unique virulence factor combinations and genomic content compared to global reference ETEC isolates
Phylogenetic relationships among Staphylococcus species and refinement of cluster groups based on multilocus data
Background: Estimates of relationships among Staphylococcus species have been hampered by poor and inconsistent resolution of phylogenies based largely on single gene analyses incorporating only a limited taxon sample. As such, the evolutionary relationships and hierarchical classification schemes among species have not been confidently established. Here, we address these points through analyses of DNA sequence data from multiple loci (16S rRNA gene, dnaJ, rpoB, and tuf gene fragments) using multiple Bayesian and maximum likelihood phylogenetic approaches that incorporate nearly all recognized Staphylococcus taxa. Results: We estimated the phylogeny of fifty-seven Staphylococcus taxa using partitioned-model Bayesian and maximum likelihood analysis, as well as Bayesian gene-tree species-tree methods. Regardless of methodology, we found broad agreement among methods that the current cluster groups require revision, although there was some disagreement among methods in resolution of higher order relationships. Based on our phylogenetic estimates, we propose a refined classification for Staphylococcus with species being classified into 15 cluster groups (based on molecular data) that adhere to six species groups (based on phenotypic properties). Conclusions: Our findings are in general agreement with gene tree-based reports of the staphylococcal phylogeny, although we identify multiple previously unreported relationships among species. Our results support the general importance of such multilocus assessments as a standard in microbial studies to more robustly infer relationships among recognized and newly discovered lineages
INVESTIGATING INVASION IN DUCTAL CARCINOMA IN SITU WITH TOPOGRAPHICAL SINGLE CELL GENOME SEQUENCING
Synchronous Ductal Carcinoma in situ (DCIS-IDC) is an early stage breast cancer invasion in which it is possible to delineate genomic evolution during invasion because of the presence of both in situ and invasive regions within the same sample. While laser capture microdissection studies of DCIS-IDC examined the relationship between the paired in situ (DCIS) and invasive (IDC) regions, these studies were either confounded by bulk tissue or limited to a small set of genes or markers. To overcome these challenges, we developed Topographic Single Cell Sequencing (TSCS), which combines laser-catapulting with single cell DNA sequencing to measure genomic copy number profiles from single tumor cells while preserving their spatial context. We applied TSCS to sequence 1,293 single cells from 10 synchronous DCIS patients. We also applied deep-exome sequencing to the in situ, invasive and normal tissues for the DCIS-IDC patients. Previous bulk tissue studies had produced several conflicting models of tumor evolution. Our data support a multiclonal invasion model, in which genome evolution occurs within the ducts and gives rise to multiple subclones that escape the ducts into the adjacent tissues to establish the invasive carcinomas. In summary, we have developed a novel method for single cell DNA sequencing, which preserves spatial context, and applied this method to understand clonal evolution during the transition between carcinoma in situ to invasive ductal carcinoma
Rebooting the human mitochondrial phylogeny: an automated and scalable methodology with expert knowledge
<p>Abstract</p> <p>Background</p> <p>Mitochondrial DNA is an ideal source of information to conduct evolutionary and phylogenetic studies due to its extraordinary properties and abundance. Many insights can be gained from these, including but not limited to screening genetic variation to identify potentially deleterious mutations. However, such advances require efficient solutions to very difficult computational problems, a need that is hampered by the very plenty of data that confers strength to the analysis.</p> <p>Results</p> <p>We develop a systematic, automated methodology to overcome these difficulties, building from readily available, public sequence databases to high-quality alignments and phylogenetic trees. Within each stage in an autonomous workflow, outputs are carefully evaluated and outlier detection rules defined to integrate expert knowledge and automated curation, hence avoiding the manual bottleneck found in past approaches to the problem. Using these techniques, we have performed exhaustive updates to the human mitochondrial phylogeny, illustrating the power and computational scalability of our approach, and we have conducted some initial analyses on the resulting phylogenies.</p> <p>Conclusions</p> <p>The problem at hand demands careful definition of inputs and adequate algorithmic treatment for its solutions to be realistic and useful. It is possible to define formal rules to address the former requirement by refining inputs directly and through their combination as outputs, and the latter are also of help to ascertain the performance of chosen algorithms. Rules can exploit known or inferred properties of datasets to simplify inputs through partitioning, therefore cutting computational costs and affording work on rapidly growing, otherwise intractable datasets. Although expert guidance may be necessary to assist the learning process, low-risk results can be fully automated and have proved themselves convenient and valuable.</p
Recommended from our members
The role of gene fusions in the evolution of metabolic pathways: the histidine biosynthesis case.
BACKGROUND: Histidine biosynthesis is one of the best characterized anabolic pathways. There is a large body of genetic and biochemical information available, including operon structure, gene expression, and increasingly larger sequence databases. For over forty years this pathway has been the subject of extensive studies, mainly in Escherichia coli and Salmonella enterica, in both of which details of histidine biosynthesis appear to be identical. In these two enterobacteria the pathway is unbranched, includes a number of unusual reactions, and consists of nine intermediates; his genes are arranged in a compact operon (hisGDC [NB]HAF [IE]), with three of them (hisNB, hisD and hisIE) coding for bifunctional enzymes. We performed a detailed analysis of his gene fusions in available genomes to understand the role of gene fusions in shaping this pathway. RESULTS: The analysis of HisA structures revealed that several gene elongation events are at the root of this protein family: internal duplication have been identified by structural superposition of the modules composing the TIM-barrel protein. Several his gene fusions happened in distinct taxonomic lineages; hisNB originated within gamma-proteobacteria and after its appearance it was transferred to Campylobacter species (epsilon-proteobacteria) and to some Bacteria belonging to the CFB group. The transfer involved the entire his operon. The hisIE gene fusion was found in several taxonomic lineages and our results suggest that it probably happened several times in distinct lineages. Gene fusions involving hisIE and hisD genes (HIS4) and hisH and hisF genes (HIS7) took place in the Eukarya domain; the latter has been transferred to some delta-proteobacteria. CONCLUSION: Gene duplication is the most widely known mechanism responsible for the origin and evolution of metabolic pathways; however, several other mechanisms might concur in the process of pathway assembly and gene fusion appeared to be one of the most important and common.RIGHTS : This article is licensed under the BioMed Central licence at http://www.biomedcentral.com/about/license which is similar to the 'Creative Commons Attribution Licence'. In brief you may : copy, distribute, and display the work; make derivative works; or make commercial use of the work - under the following conditions: the original author must be given credit; for any reuse or distribution, it must be made clear to others what the license terms of this work are
Confronting Sources of Systematic Error to Resolve Historically Contentious Relationships: A Case Study Using Gadiform Fishes (Teleostei, Paracanthopterygii, Gadiformes)
Reliable estimation of phylogeny is central to avoid inaccuracy in downstream macroevolutionary inferences. However, limitations exist in the implementation of concatenated and summary coalescent approaches, and Bayesian and full coalescent inference methods may not yet be feasible for computation of phylogeny using complicated models and large data sets. Here, we explored methodological (e.g., optimality criteria, character sampling, model selection) and biological (e.g., heterotachy, branch length heterogeneity) sources of systematic error that can result in biased or incorrect parameter estimates when reconstructing phylogeny by using the gadiform fishes as a model clade. Gadiformes include some of the most economically important fishes in the world (e.g., Cods, Hakes, and Rattails). Despite many attempts, a robust higher-level phylogenetic framework was lacking due to limited character and taxonomic sampling, particularly from several species-poor families that have been recalcitrant to phylogenetic placement. We compiled the first phylogenomic data set, including 14,208 loci (\u3e role= presentation \u3e\u3e2.8 M bp) from 58 species representing all recognized gadiform families, to infer a time-calibrated phylogeny for the group. Data were generated with a gene-capture approach targeting coding DNA sequences from single-copy protein-coding genes. Species-tree and concatenated maximum-likelihood (ML) analyses resolved all family-level relationships within Gadiformes. While there were a few differences between topologies produced by the DNA and the amino acid data sets, most of the historically unresolved relationships among gadiform lineages were consistently well resolved with high support in our analyses regardless of the methodological and biological approaches used. However, at deeper levels, we observed inconsistency in branch support estimates between bootstrap and gene and site coefficient factors (gCF, sCF). Despite numerous short internodes, all relationships received unequivocal bootstrap support while gCF and sCF had very little support, reflecting hidden conflict across loci. Most of the gene-tree and species-tree discordance in our study is a result of short divergence times, and consequent lack of informative characters at deep levels, rather than incomplete lineage sorting. We use this phylogeny to establish a new higher-level classification of Gadiformes as a way of clarifying the evolutionary diversification of the order. We recognize 17 families in five suborders: Bregmacerotoidei, Gadoidei, Ranicipitoidei, Merluccioidei, and Macrouroidei (including two subclades). A time-calibrated analysis using 15 fossil taxa suggests that Gadiformes evolved ∼ role= presentation \u3e∼79.5 Ma in the late Cretaceous, but that most extant lineages diverged after the Cretaceous–Paleogene (K-Pg) mass extinction (66 Ma). Our results reiterate the importance of examining phylogenomic analyses for evidence of systematic error that can emerge as a result of unsuitable modeling of biological factors and/or methodological issues, even when data sets are large and yield high support for phylogenetic relationships. [Branch length heterogeneity; Codfishes; commercial fish species; Cretaceous-Paleogene (K-Pg); heterotachy; systematic error; target enrichment.
- …