Search CORE

1,767 research outputs found

CERN openlab Whitepaper on Future IT Challenges in Scientific Research

Author: Di Meglio Alberto
Gaillard Melissa
Purcell Andrew
Publication venue
Publication date: 01/01/2014
Field of study

This whitepaper describes the major IT challenges in scientific research at CERN and several other European and international research laboratories and projects. Each challenge is exemplified through a set of concrete use cases drawn from the requirements of large-scale scientific programs. The paper is based on contributions from many researchers and IT experts of the participating laboratories and also input from the existing CERN openlab industrial sponsors. The views expressed in this document are those of the individual contributors and do not necessarily reflect the view of their organisations and/or affiliates

ZENODO

CERN Document Server

Efficient, Dependable Storage of Human Genome Sequencing Data

Author: Cogo Vinicius
Publication venue
Publication date: 01/05/2020
Field of study

A compreensão do genoma humano impacta várias áreas da vida. Os dados oriundos do genoma humano são enormes pois existem milhões de amostras a espera de serem sequenciadas e cada genoma humano sequenciado pode ocupar centenas de gigabytes de espaço de armazenamento. Os genomas humanos são críticos porque são extremamente valiosos para a investigação e porque podem fornecer informações delicadas sobre o estado de saúde dos indivíduos, identificar os seus dadores ou até mesmo revelar informações sobre os parentes destes. O tamanho e a criticidade destes genomas, para além da quantidade de dados produzidos por instituições médicas e de ciências da vida, exigem que os sistemas informáticos sejam escaláveis, ao mesmo tempo que sejam seguros, confiáveis, auditáveis e com custos acessíveis. As infraestruturas de armazenamento existentes são tão caras que não nos permitem ignorar a eficiência de custos no armazenamento de genomas humanos, assim como em geral estas não possuem o conhecimento e os mecanismos adequados para proteger a privacidade dos dadores de amostras biológicas. Esta tese propõe um sistema de armazenamento de genomas humanos eficiente, seguro e auditável para instituições médicas e de ciências da vida. Ele aprimora os ecossistemas de armazenamento tradicionais com técnicas de privacidade, redução do tamanho dos dados e auditabilidade a fim de permitir o uso eficiente e confiável de infraestruturas públicas de computação em nuvem para armazenar genomas humanos. As contribuições desta tese incluem (1) um estudo sobre a sensibilidade à privacidade dos genomas humanos; (2) um método para detetar sistematicamente as porções dos genomas que são sensíveis à privacidade; (3) algoritmos de redução do tamanho de dados, especializados para dados de genomas sequenciados; (4) um esquema de auditoria independente para armazenamento disperso e seguro de dados; e (5) um fluxo de armazenamento completo que obtém garantias razoáveis de proteção, segurança e confiabilidade a custos modestos (por exemplo, menos de

1/Genoma/Ano), integrando os mecanismos propostos a configurações de armazenamento apropriadasThe understanding of human genome impacts several areas of human life. Data from human genomes is massive because there are millions of samples to be sequenced, and each sequenced human genome may size hundreds of gigabytes. Human genomes are critical because they are extremely valuable to research and may provide hints on individuals’ health status, identify their donors, or reveal information about donors’ relatives. Their size and criticality, plus the amount of data being produced by medical and life-sciences institutions, require systems to scale while being secure, dependable, auditable, and affordable. Current storage infrastructures are too expensive to ignore cost efficiency in storing human genomes, and they lack the proper knowledge and mechanisms to protect the privacy of sample donors. This thesis proposes an efficient storage system for human genomes that medical and lifesciences institutions may trust and afford. It enhances traditional storage ecosystems with privacy-aware, data-reduction, and auditability techniques to enable the efficient, dependable use of multi-tenant infrastructures to store human genomes. Contributions from this thesis include (1) a study on the privacy-sensitivity of human genomes; (2) to detect genomes’ privacy-sensitive portions systematically; (3) specialised data reduction algorithms for sequencing data; (4) an independent auditability scheme for secure dispersed storage; and (5) a complete storage pipeline that obtains reasonable privacy protection, security, and dependability guarantees at modest costs (e.g., less than

1/Genome/Year) by integrating the proposed mechanisms with appropriate storage configurations

Universidade de Lisboa: Repositório.UL

‘Next-Generation’ surveillance: an epidemiologists’ perspective on the use of molecular information in food safety and animal health decision-making

Author: Dufour S.
Muellner P.
Stärk K.D.C.
Zadoks R.N.
Publication venue: 'Wiley'
Publication date: 01/08/2016
Field of study

Advances in the availability and affordability of molecular and genomic data are transforming human health care. Surveillance aimed at supporting and improving food safety and animal health is likely to undergo a similar transformation. We propose a definition of ‘molecular surveillance’ in this context and argue that molecular data are an adjunct to rather than a substitute for sound epidemiological study and surveillance design. Specific considerations with regard to sample collection are raised, as is the importance of the relation between the molecular clock speed of genetic markers and the spatiotemporal scale of the surveillance activity, which can be control- or strategy-focused. Development of standards for study design and assessment of molecular surveillance system attributes is needed, together with development of an interdisciplinary skills base covering both molecular and epidemiological principles

Enlighten

Machine Learning and Integrative Analysis of Biomedical Big Data.

Author: Choi Howard
Chung Neo Christopher
Mirza Bilal
Ping Peipei
Wang Jie
Wang Wei
Publication venue: eScholarship, University of California
Publication date: 01/01/2019
Field of study

Recent developments in high-throughput technologies have accelerated the accumulation of massive amounts of omics data from multiple sources: genome, epigenome, transcriptome, proteome, metabolome, etc. Traditionally, data from each source (e.g., genome) is analyzed in isolation using statistical and machine learning (ML) methods. Integrative analysis of multi-omics and clinical data is key to new biomedical discoveries and advancements in precision medicine. However, data integration poses new computational challenges as well as exacerbates the ones associated with single-omics studies. Specialized computational approaches are required to effectively and efficiently perform integrative analysis of biomedical data acquired from diverse modalities. In this review, we discuss state-of-the-art ML-based approaches for tackling five specific computational challenges associated with integrative analysis: curse of dimensionality, data heterogeneity, missing data, class imbalance and scalability issues

Multidisciplinary Digital Publishing Institute

Ezid

Directory of Open Access Journals

eScholarship - University of California

Boosting analyses in the life sciences via clusters, grids and clouds

Author: Carretero Pérez Jesús
García Blas Francisco Javier
Gesing Sandra
Montagnat Johan
Publication venue: 'Elsevier BV'
Publication date: 01/02/2017
Field of study

In the last 20 years, computational methods have become an important part of developing emerging technologies for the field of bioinformatics and biomedicine. Those methods rely heavily on large scale computational resources as they need to manage Tbytes or Pbytes of data with large-scale structural and functional relationships, TFlops or PFlops of computing power for simulating highly complex models, or many-task processes and workflows for processing and analyzing data. This special issue contains papers showing existing solutions and latest developments in Life Sciences and Computing Sciences to collaboratively explore new ideas and approaches to successfully apply distributed IT-systems in translational research, clinical intervention, and decision-making. (C) 2016 Published by Elsevier B.V

Universidad Carlos III de Madrid e-Archivo

Towards a European Health Research and Innovation Cloud (HRIC)

Author: Aarestrup Frank
Albeyatti Abdullah
Armitage Willian
Auffray C
Augello Luca
Balling Rudi
Benhabiles Nora
Bertolini Guido
Bjaalie Jan
Black Michaela
Blomberg Niklas
Bogaert Petronille
Bubak Marian
Claerhout Brecht
Clarke Laura
D'Errico Gianni
De Meulder Bertrand
Di Meglio Alberto
Forgo Nikolaus
Gans-Combe Caroline
Gray Alexander
Gut Ivo
Gyllenberg Alexandra
Hemmrich-Stanisak Georg
Hjorth Lars
Ioannidis Yannis
Jarmalaite Sonata
Kel Alexander
Kherif Ferath
Korbel Jan
Larue Catherine
László Mitzi
Maas Andrew
Magalhaes Luis
Manneh-Vangramberen Isabelle
Morley-Fletcher Edwin
Ohmann Christian
Oksvold Per
Oxtoby Neil
Perseil Isabelle
Pezoulas Vasileios
Riess Olaf
Riper Heleen
Roca Josep
Rosenstiel Philip
Sabatier Philippe
Sanz Ferran
Tayeb Mohammed
Thomassen Gard
Van Bussel Johan
Van Den Bulcke Marc
Van Oyen Herman
Publication venue
Publication date: 01/01/2020
Field of study

The European Union (EU) initiative on the Digital Transformation of Health and Care (Digicare) aims to provide the conditions necessary for building a secure, flexible, and decentralized digital health infrastructure. Creating a European Health Research and Innovation Cloud (HRIC) within this environment should enable data sharing and analysis for health research across the EU, in compliance with data protection legislation while preserving the full trust of the participants. Such a HRIC should learn from and build on existing data infrastructures, integrate best practices, and focus on the concrete needs of the community in terms of technologies, governance, management, regulation, and ethics requirements. Here, we describe the vision and expected benefits of digital data sharing in health research activities and present a roadmap that fosters the opportunities while answering the challenges of implementing a HRIC. For this, we put forward five specific recommendations and action points to ensure that a European HRIC: i) is built on established standards and guidelines, providing cloud technologies through an open and decentralized infrastructure; ii) is developed and certified to the highest standards of interoperability and data security that can be trusted by all stakeholders; iii) is supported by a robust ethical and legal framework that is compliant with the EU General Data Protection Regulation (GDPR); iv) establishes a proper environment for the training of new generations of data and medical scientists; and v) stimulates research and innovation in transnational collaborations through public and private initiatives and partnerships funded by the EU through Horizon 2020 and Horizon Europe

VU Research Portal

ZENODO

Serveur académique lausannois

Publikationsserver der Universität Tübingen

Ulster University's Research Portal

UPF Digital Repository

NORA - Norwegian Open Research Archives

Online Research Database In Technology

Lund University Publications

Sciensano Publications Repository

Ghent University Academic Bibliography

UCL Discovery