36 research outputs found

    Developing a manually annotated clinical document corpus to identify phenotypic information for inflammatory bowel disease

    Get PDF
    <p>Abstract</p> <p>Background</p> <p>Natural Language Processing (NLP) systems can be used for specific Information Extraction (IE) tasks such as extracting phenotypic data from the electronic medical record (EMR). These data are useful for translational research and are often found only in free text clinical notes. A key required step for IE is the manual annotation of clinical corpora and the creation of a reference standard for (1) training and validation tasks and (2) to focus and clarify NLP system requirements. These tasks are time consuming, expensive, and require considerable effort on the part of human reviewers.</p> <p>Methods</p> <p>Using a set of clinical documents from the VA EMR for a particular use case of interest we identify specific challenges and present several opportunities for annotation tasks. We demonstrate specific methods using an open source annotation tool, a customized annotation schema, and a corpus of clinical documents for patients known to have a diagnosis of Inflammatory Bowel Disease (IBD). We report clinician annotator agreement at the document, concept, and concept attribute level. We estimate concept yield in terms of annotated concepts within specific note sections and document types.</p> <p>Results</p> <p>Annotator agreement at the document level for documents that contained concepts of interest for IBD using estimated Kappa statistic (95% CI) was very high at 0.87 (0.82, 0.93). At the concept level, F-measure ranged from 0.61 to 0.83. However, agreement varied greatly at the specific concept attribute level. For this particular use case (IBD), clinical documents producing the highest concept yield per document included GI clinic notes and primary care notes. Within the various types of notes, the highest concept yield was in sections representing patient assessment and history of presenting illness. Ancillary service documents and family history and plan note sections produced the lowest concept yield.</p> <p>Conclusion</p> <p>Challenges include defining and building appropriate annotation schemas, adequately training clinician annotators, and determining the appropriate level of information to be annotated. Opportunities include narrowing the focus of information extraction to use case specific note types and sections, especially in cases where NLP systems will be used to extract information from large repositories of electronic clinical note documents.</p

    Biomedical informatics and translational medicine

    Get PDF
    Biomedical informatics involves a core set of methodologies that can provide a foundation for crossing the "translational barriers" associated with translational medicine. To this end, the fundamental aspects of biomedical informatics (e.g., bioinformatics, imaging informatics, clinical informatics, and public health informatics) may be essential in helping improve the ability to bring basic research findings to the bedside, evaluate the efficacy of interventions across communities, and enable the assessment of the eventual impact of translational medicine innovations on health policies. Here, a brief description is provided for a selection of key biomedical informatics topics (Decision Support, Natural Language Processing, Standards, Information Retrieval, and Electronic Health Records) and their relevance to translational medicine. Based on contributions and advancements in each of these topic areas, the article proposes that biomedical informatics practitioners ("biomedical informaticians") can be essential members of translational medicine teams

    A Genome-Wide Association Study of Diabetic Kidney Disease in Subjects With Type 2 Diabetes

    Get PDF
    dentification of sequence variants robustly associated with predisposition to diabetic kidney disease (DKD) has the potential to provide insights into the pathophysiological mechanisms responsible. We conducted a genome-wide association study (GWAS) of DKD in type 2 diabetes (T2D) using eight complementary dichotomous and quantitative DKD phenotypes: the principal dichotomous analysis involved 5,717 T2D subjects, 3,345 with DKD. Promising association signals were evaluated in up to 26,827 subjects with T2D (12,710 with DKD). A combined T1D+T2D GWAS was performed using complementary data available for subjects with T1D, which, with replication samples, involved up to 40,340 subjects with diabetes (18,582 with DKD). Analysis of specific DKD phenotypes identified a novel signal near GABRR1 (rs9942471, P = 4.5 x 10(-8)) associated with microalbuminuria in European T2D case subjects. However, no replication of this signal was observed in Asian subjects with T2D or in the equivalent T1D analysis. There was only limited support, in this substantially enlarged analysis, for association at previously reported DKD signals, except for those at UMOD and PRKAG2, both associated with estimated glomerular filtration rate. We conclude that, despite challenges in addressing phenotypic heterogeneity, access to increased sample sizes will continue to provide more robust inference regarding risk variant discovery for DKD.Peer reviewe
    corecore