Search CORE

Columbia University Academic Commons

Directory of Open Access Journals

FigShare

STORMSeq: An Open-Source, User-Friendly Pipeline for Processing Personal Genomics Data in the Cloud

Author: Dudley Joel T.
Fernald Guy Haskin
Karczewski Konrad J.
Martin Alicia R.
Snyder Michael
Tatonetti Nicholas P.
Publication venue: 'Columbia University Libraries/Information Services'
Publication date: 01/01/2014
Field of study

The increasing public availability of personal complete genome sequencing data has ushered in an era of democratized genomics. However, read mapping and variant calling software is constantly improving and individuals with personal genomic data may prefer to customize and update their variant calls. Here, we describe STORMSeq (Scalable Tools for Open-Source Read Mapping), a graphical interface cloud computing solution that does not require a parallel computing environment or extensive technical experience. This customizable and modular system performs read mapping, read cleaning, and variant calling and annotation. At present, STORMSeq costs approximately

2 and 5–10 hours to process a full exome sequence and

30 and 3–8 days to process a whole genome sequence. We provide this open-access and open-source resource as a user-friendly interface in Amazon EC2

Columbia University Academic Commons

Directory of Open Access Journals

Infoscience - École polytechnique fédérale de Lausanne

FigShare

Quantifying supercoiling-induced denaturation bubbles in DNA

Author: Adamcik Jozef
Dietler Giovanni
Jeon Jae-Hyung
Karczewski Konrad J.
Metzler Ralf
Publication venue: 'Royal Society of Chemistry (RSC)'
Publication date: 27/02/2013
Field of study

In both eukaryotic and prokaryotic DNA sequences of 30-100 base-pairs rich in AT base-pairs have been identified at which the double helix preferentially unwinds. Such DNA unwinding elements are commonly associated with origins for DNA replication and transcription, and with chromosomal matrix attachment regions. Here we present a quantitative study of local DNA unwinding based on extensive single DNA plasmid imaging. We demonstrate that long-lived single-stranded denaturation bubbles exist in negatively supercoiled DNA, at the expense of partial twist release. Remarkably, we observe a linear relation between the degree of supercoiling and the bubble size, in excellent agreement with statistical modelling. Furthermore, we obtain the full distribution of bubble sizes and the opening probabilities at varying salt and temperature conditions. The results presented herein underline the important role of denaturation bubbles in negatively supercoiled DNA for biological processes such as transcription and replication initiation in vivo

Quantifying unobserved protein-coding variants in human populations provides a roadmap for large-scale sequencing projects

Author: Daniel MacArthur
Gregory Valiant
James Zou
Kaitlin Samocha
Konrad Karczewski
Mark Daly
Monkol Lek
null null
Paul Valiant
Shamil Sunyaev
Siu On Chan
Publication venue: 'Springer Science and Business Media LLC'
Publication date: 06/11/2015
Field of study

As new proposals aim to sequence ever larger collection of humans, it is critical to have a quantitative framework to evaluate the statistical power of these projects. We developed a new algorithm, UnseenEst, and applied it to the exomes of 60,706 individuals to estimate the frequency distribution of all protein-coding variants, including rare variants that have not been observed yet in the current cohorts. Our results quantified the number of new variants that we expect to identify as sequencing cohorts reach hundreds of thousands of individuals. With 500K individuals, we find that we expect to capture 7.5% of all possible loss-of-function variants and 12% of all possible missense variants. We also estimate that 2,900 genes have loss-of-function frequency of <0.00001 in healthy humans, consistent with very strong intolerance to gene inactivation.United States. National Institutes of Health (U54DK105566)United States. National Institutes of Health (R01GM104371

DSpace@MIT

Harvard University - DASH

An experience report on (auto-)tuning of mesh-based PDE solvers on shared memory systems.

Author: Charrier Dominic E.
Deelman Ewa
Dongarra J.J.
Karczewski Konrad
Weinzierl Tobias
Wyrzykowski Roman
Publication venue: 'Springer Science and Business Media LLC'
Publication date: 01/01/2018
Field of study

With the advent of manycore systems, shared memory parallelisation has gained importance in high performance computing. Once a code is decomposed into tasks or parallel regions, it becomes crucial to identify reasonable grain sizes, i.e. minimum problem sizes per task that make the algorithm expose a high concurrency at low overhead. Many papers do not detail what reasonable task sizes are, and consider their findings craftsmanship not worth discussion. We have implemented an autotuning algorithm, a machine learning approach, for a project developing a hyperbolic equation system solver. Autotuning here is important as the grid and task workload are multifaceted and change frequently during runtime. In this paper, we summarise our lessons learned. We infer tweaks and idioms for general autotuning algorithms and we clarify that such a approach does not free users completely from grain size awareness

Durham Research Online

SAIGE-GENE plus improves the efficiency and accuracy of set-based rare variant association tests

Author: Bi Wenjian
Daly Mark J.
Dey Kushal K.
Jagadeesh Karthik A.
Karczewski Konrad J.
Lee Seunggeun
Neale Benjamin M.
Zhao Zhangchen
Zhou Wei
Publication venue
Publication date: 01/01/2022
Field of study

Several biobanks, including UK Biobank (UKBB), are generating large-scale sequencing data. An existing method, SAIGE-GENE, performs well when testing variants with minor allele frequency (MAF) SAIGE-GENE+ performs set-based rare variant association tests with improved type 1 error control and computational efficiency by collapsing ultra-rare variants and conducting multiple tests corresponding to different minor allele frequency cutoffs and annotations.Peer reviewe

Fast DEM collision checks on multicore nodes.

Author: Deelman Ewa
Dongarra J.J.
Karczewski Konrad
Koziara Tomasz
Krestenitis Konstantinos
Weinzierl Tobias
Wyrzykowski Roman
Publication venue: 'Springer Science and Business Media LLC'
Publication date: 01/01/2018
Field of study

Many particle simulations today rely on spherical or analytical particle shape descriptions. They find non-spherical, triangulated particle models computationally infeasible due to expensive collision detections. We propose a hybrid collision detection algorithm based upon an iterative solve of a minimisation problem that automatically falls back to a brute-force comparison-based algorithm variant if the problem is ill-posed. Such a hybrid can exploit the vector facilities of modern chips and it is well-prepared for the arising manycore era. Our approach pushes the boundary where non-analytical particle shapes and the aligning of more accurate first principle physics become manageable

Durham Research Online

Base-specific mutational intolerance near splice sites clarifies the role of nonessential splice nucleotides

Author: Daly Emma
Daly Mark J.
Karczewski Konrad J.
MacArthur Daniel G.
Neale Benjamin M.
Rivas Manuel A.
Samocha Kaitlin E.
Schmandt Ben
Zhang Sidi
Publication venue
Publication date: 01/06/2018
Field of study

Variation in RNA splicing (i.e., alternative splicing) plays an important role in many diseases. Variants near 5' and 3' splice sites often affect splicing, but the effects of these variants on splicing and disease have not been fully characterized beyond the two "essential" splice nucleotides flanking each exon. Here we provide quantitative measurements of tolerance to mutational disruptions by position and reference allele-alternative allele combinations. We show that certain reference alleles are particularly sensitive to mutations, regardless of the alternative alleles into which they are mutated. Using public RNA-seq data, we demonstrate that individuals carrying such variants have significantly lower levels of the correctly spliced transcript, compared to individuals without them, and confirm that these specific substitutions are highly enriched for known Mendelian mutations. Our results propose a more refined definition of the "splice region" and offer a new way to prioritize and provide functional interpretation of variants identified in diagnostic sequencing and association studies.Peer reviewe

A structural variation reference for medical and population genetics

Author: Brand Harrison
Collins Ryan L.
Färkkilä Martti
Genome Aggregation Database Consor
Genome Aggregation Database Prod T
Groop Leif
Holi Matti M.
Kallela Mikko
Kaprio Jaakko
Karczewski Konrad J.
Palotie Aarno
Ripatti Samuli
Talkowski Michael E.
Tuomi Tiinamaija
Wessman Maija
Publication venue
Publication date: 28/05/2020
Field of study

Structural variants (SVs) rearrange large segments of DNA(1) and can have profound consequences in evolution and human disease(2,3). As national biobanks, disease-association studies, and clinical genetic testing have grown increasingly reliant on genome sequencing, population references such as the Genome Aggregation Database (gnomAD)(4) have become integral in the interpretation of single-nucleotide variants (SNVs)(5). However, there are no reference maps of SVs from high-coverage genome sequencing comparable to those for SNVs. Here we present a reference of sequence-resolved SVs constructed from 14,891 genomes across diverse global populations (54% non-European) in gnomAD. We discovered a rich and complex landscape of 433,371 SVs, from which we estimate that SVs are responsible for 25-29% of all rare protein-truncating events per genome. We found strong correlations between natural selection against damaging SNVs and rare SVs that disrupt or duplicate protein-coding sequence, which suggests that genes that are highly intolerant to loss-of-function are also sensitive to increased dosage(6). We also uncovered modest selection against noncoding SVs in cis-regulatory elements, although selection against protein-truncating SVs was stronger than all noncoding effects. Finally, we identified very large (over one megabase), rare SVs in 3.9% of samples, and estimate that 0.13% of individuals may carry an SV that meets the existing criteria for clinically important incidental findings(7). This SV resource is freely distributed via the gnomAD browser(8) and will have broad utility in population genetics, disease-association studies, and diagnostic screening.Peer reviewe

Evaluating drug targets through human loss-of-function genetic variation

Author: Daly Mark J.
Färkkilä Martti
Genome Aggregation Database Consor
Genome Aggregation Database Prod T
Groop Leif
Holi Matti M.
Kallela Mikko
Kaprio Jaakko
Karczewski Konrad J.
MacArthur Daniel G.
Martin Hilary C.
Minikel Eric Vallabh
Palotie Aarno
Ripatti Samuli
Tuomi Tiinamaija
Wessman Maija
Publication venue
Publication date: 28/05/2020
Field of study

Naturally occurring human genetic variants that are predicted to inactivate protein-coding genes provide an in vivo model of human gene inactivation that complements knockout studies in cells and model organisms. Here we report three key findings regarding the assessment of candidate drug targets using human loss-of-function variants. First, even essential genes, in which loss-of-function variants are not tolerated, can be highly successful as targets of inhibitory drugs. Second, in most genes, loss-of-function variants are sufficiently rare that genotype-based ascertainment of homozygous or compound heterozygous 'knockout' humans will await sample sizes that are approximately 1,000 times those presently available, unless recruitment focuses on consanguineous individuals. Third, automated variant annotation and filtering are powerful, but manual curation remains crucial for removing artefacts, and is a prerequisite for recall-by-genotype efforts. Our results provide a roadmap for human knockout studies and should guide the interpretation of loss-of-function variants in drug development.Peer reviewe