Search CORE

2 research outputs found

Inverting the model of genomics data sharing with the NHGRI Genomic Data Science Analysis, Visualization, and Informatics Lab-space

Author: Abel Haley J
Afgan Enis
Baker Dannon
Banasiewicz M Katie
Banks Eric
Baumann Alexander
Baumann Michael
Bernard Clare
Blauvelt Lon
Cabansay Louise
Caetano-Anollés Derek
Canas Justin
Carey Vincent J
Carroll Robert J
Chaluvadi Sushma
Chilton John
Clements Dave
Cox Katherine EL
Culotti Alessandro
Di Francesco Valentina
Disman William
Ellrott Kyle
Geistlinger Ludwig
Ghanaim Elena M
Goecks Jeremy
Golitsynskiy Sergey
Grossman Robert L
Gupta Namrata
Hajian Allie
Hall Ira M
Hannafious Brian
Hansen Kasper D
Harris Tim
Hastie Mim
Herman Kate
Hutter Carolyn
Jalili Vahid
Kammers Kai
Kiernan Elizabeth
Kovalsy Anton
Kucher Nataliya
Lawson Jonathan
Leek Jeffrey T
Lucas Julian
Luria Anne O’Donnell
Mahmoud Alexandru
McDade Frances
Morgan Martin
Mosher Stephen
Munshi Ruchi
Nekrutenko Anton
Oh Sehyun
Osborn Kevin
Ostrovsky Alexander
Overbeck Charles
O’Connor Brian D
O’Farrell Ash
Paten Benedict
Patterson Candace
Philippakis Anthony A
Ramos Marcel
Reddy Radhika
Reeves Valerie
Reid Charles
Rogers Dave
Rula Andrew
s Yuen Deni
Sargent Luke
Schatz Michael C
Sen Shurjo K
Sheets Elizabeth A
Shepherd Lori
Simeon Marianie
Steinberg David Charles
Stevens Ana
Stubbs BJ
Suderman Keith
Tan Frederick J
Taylor Casey Overby
Taylor M Morgan
Thomas Salin
Title Robert
Torstenson Eric
Turaga Nitesh
Van der Auwera Geraldine A
Vessio Jennifer
Vizzier Benton A
Vosburg Trish
Waldron Levi
Walker Jason
Walsh Brian
Wang Qi
Wang Ting
Warren Noah
Wellington Christopher
Wheelan Sarah J
Wiley Ken L
Wuichet Kristin
Yuksel Kaan
Zarate Samantha
Publication venue: 'Elsevier BV'
Publication date: 12/01/2022
Field of study

The NHGRI Genomic Data Science Analysis, Visualization, and Informatics Lab-space (AnVIL; https://anvilproject.org) was developed to address a widespread community need for a unified computing environment for genomics data storage, management, and analysis. In this perspective, we present AnVIL, describe its ecosystem and interoperability with other platforms, and highlight how this platform and associated initiatives contribute to improved genomic data sharing efforts. The AnVIL is a federated cloud platform designed to manage and store genomics and related data, enable population-scale analysis, and facilitate collaboration through the sharing of data, code, and analysis results. By inverting the traditional model of data sharing, the AnVIL eliminates the need for data movement while also adding security measures for active threat detection and monitoring and provides scalable, shared computing resources for any researcher. We describe the core data management and analysis components of the AnVIL, which currently consists of Terra, Gen3, Galaxy, RStudio/Bioconductor, Dockstore, and Jupyter, and describe several flagship genomics datasets available within the AnVIL. We continue to extend and innovate the AnVIL ecosystem by implementing new capabilities, including mechanisms for interoperability and responsible data sharing, while streamlining access management. The AnVIL opens many new opportunities for analysis, collaboration, and data sharing that are needed to drive research and to make discoveries through the joint analysis of hundreds of thousands to millions of genomes along with associated clinical and molecular data types

Cold Spring Harbor Laboratory Institutional Repository

Diversifying the genomic data science research community

Over the past 20 years, the explosion of genomic data collection and the cloud computing revolution have made computational and data science research accessible to anyone with a web browser and an internet connection. However, students at institutions with limited resources have received relatively little exposure to curricula or professional development opportunities that lead to careers in genomic data science. To broaden participation in genomics research, the scientific community needs to support these programs in local education and research at underserved institutions (UIs). These include community colleges, historically Black colleges and universities, Hispanic-serving institutions, and tribal colleges and universities that support ethnically, racially, and socioeconomically underrepresented students in the United States. We have formed the Genomic Data Science Community Network to support students, faculty, and their networks to identify opportunities and broaden access to genomic data science. These opportunities include expanding access to infrastructure and data, providing UI faculty development opportunities, strengthening collaborations among faculty, recognizing UI teaching and research excellence, fostering student awareness, developing modular and open-source resources, expanding course-based undergraduate research experiences (CUREs), building curriculum, supporting student professional development and research, and removing financial barriers through funding programs and collaborator support

arXiv.org e-Print Archive

eScholarship - University of California