An annotated corpus with nanomedicine and pharmacokinetic parameters

Jimenez, Ivan; Lewinski, Nastassja; McInnes, Bridget

research

An annotated corpus with nanomedicine and pharmacokinetic parameters

Authors: Ivan Jimenez
Nastassja Lewinski
Bridget McInnes
Publication date: 1 January 2017
Publisher: VCU Scholars Compass
Doi

Abstract

A vast amount of data on nanomedicines is being generated and published, and natural language processing (NLP) approaches can automate the extraction of unstructured text-based data. Annotated corpora are a key resource for NLP and information extraction methods which employ machine learning. Although corpora are available for pharmaceuticals, resources for nanomedicines and nanotechnology are still limited. To foster nanotechnology text mining (NanoNLP) efforts, we have constructed a corpus of annotated drug product inserts taken from the US Food and Drug Administration’s Drugs@FDA online database. In this work, we present the development of the Engineered Nanomedicine Database corpus to support the evaluation of nanomedicine entity extraction. The data were manually annotated for 21 entity mentions consisting of nanomedicine physicochemical characterization, exposure, and biologic response information of 41 Food and Drug Administration-approved nanomedicines. We evaluate the reliability of the manual annotations and demonstrate the use of the corpus by evaluating two state-of-the-art named entity extraction systems, OpenNLP and Stanford NER. The annotated corpus is available open source and, based on these results, guidelines and suggestions for future development of additional nanomedicine corpora are provided

Similar works

Full text

Open in the Core reader

Download PDF

Available Versions

VCU Scholars Compass

oai:scholarscompass.vcu.edu:cl...

Last time updated on 09/07/2019

Crossref

info:doi/10.2147%2Fijn.s137117

Last time updated on 01/04/2019