Search CORE

255 research outputs found

Near-optimal RNA-Seq quantification

Author: Bray Nicolas
Melsted Páll
Pachter Lior
Pimentel Harold
Publication venue
Publication date: 11/05/2015
Field of study

We present a novel approach to RNA-Seq quantification that is near optimal in speed and accuracy. Software implementing the approach, called kallisto, can be used to analyze 30 million unaligned paired-end RNA-Seq reads in less than 5 minutes on a standard laptop computer while providing results as accurate as those of the best existing tools. This removes a major computational bottleneck in RNA-Seq analysis.Comment: - Added some results (paralog analysis, allele specific expression analysis, alignment comparison, accuracy analysis with TPMs) - Switched bootstrap analysis to human sample from SEQC-MAQCIII - Provided link to a snakefile that allows for reproducibility of all results and figures in the pape

arXiv.org e-Print Archive

Caltech Authors

A Survey of Instrumental Music Actvity in The Small High Schools in North Dakota

Author: Melsted Haraldur Freeman
Publication venue: UND Scholarly Commons
Publication date: 01/08/1954
Field of study

UND Scholarly Commons (University of North Dakota)

Barcode, UMI, Set format and BUStools

Author: Melsted Páll
Ntranos Vasilis
Pachter Lior
Publication venue: 'Oxford University Press (OUP)'
Publication date: 01/11/2019
Field of study

We introduce the Barcode-UMI-Set format (BUS) for representing pseudoalignments of reads from single-cell RNA-seq experiments. The format can be used with all single-cell RNA-seq technologies, and we show that BUS files can be efficiently generated. BUStools is a suite of tools for working with BUS files and facilitates rapid quantification and analysis of single-cell RNA-seq data. The BUS format therefore makes possible the development of modular, technology-specific and robust workflows for single-cell RNA-seq analysis

Caltech Authors

Pseudoalignment for metagenomic read assignment

Author: Bray N.
Melsted P.
Pachter L.
Pimentel H.
Schaeffer L.
Publication venue: 'Oxford University Press (OUP)'
Publication date: 15/07/2017
Field of study

Motivation: Read assignment is an important first step in many metagenomic analysis workflows, providing the basis for identification and quantification of species. However ambiguity among the sequences of many strains makes it difficult to assign reads at the lowest level of taxonomy, and reads are typically assigned to taxonomic levels where they are unambiguous. We explore connections between metagenomic read assignment and the quantification of transcripts from RNA-Seq data in order to develop novel methods for rapid and accurate quantification of metagenomic strains. Results: We find that the recent idea of pseudoalignment introduced in the RNA-Seq context is highly applicable in the metagenomics setting. When coupled with the Expectation-Maximization (EM) algorithm, reads can be assigned far more accurately and quickly than is currently possible with state of the art software, making it possible and practical for the first time to analyze abundances of individual genomes in metagenomics projects

Caltech Authors

A discriminative learning approach to differential expression analysis for single-cell RNA-seq

Author: Melsted Páll
Ntranos Vasilis
Pachter Lior
Yi Lynn
Publication venue: Nature Publishing Group
Publication date: 01/02/2019
Field of study

Single-cell RNA-seq makes it possible to characterize the transcriptomes of cell types across different conditions and to identify their transcriptional signatures via differential analysis. Our method detects changes in transcript dynamics and in overall gene abundance in large numbers of cells to determine differential expression. When applied to transcript compatibility counts obtained via pseudoalignment, our approach provides a quantification-free analysis of 3′ single-cell RNA-seq that can identify previously undetectable marker genes

Caltech Authors

A direct comparison of genome alignment and transcriptome pseudoalignment

Author: Liu Lauren
Melsted Páll
Pachter Lior
Yi Lynn
Publication venue
Publication date: 16/10/2018
Field of study

Motivation: Genome alignment of reads is the first step of most genome analysis workflows. In the case of RNA-Seq, transcriptome pseudoalignment of reads is a fast alternative to genome alignment, but the different coordinate systems of the genome and transcriptome have made it difficult to perform direct comparisons between the approaches. Results: We have developed tools for converting genome alignments to transcriptome pseudoalignments, and conversely, for projecting transcriptome pseudoalignments to genome alignments. Using these tools, we performed a direct comparison of genome alignment with transcriptome pseudoalignment. We find that both approaches produce similar quantifications. This means that for many applications genome alignment and transcriptome pseudoalignment are interchangeable. Availability and Implementation: bam2tcc is a C++14 software for converting alignments in SAM/BAM format to transcript compatibility counts (TCCs) and is available at https://github.com/pachterlab/bam2tcc. kallisto genomebam is a user option of kallisto that outputs a sorted BAM file in genome coordinates as part of transcriptome pseudoalignment. The feature has been released with kallisto v0.44.0, and is available at https://pachterlab.github.io/kallisto/

Caltech Authors