Search CORE

1,332 research outputs found

The Research Object Suite of Ontologies: Sharing and Exchanging Research Data and Methods on the Open Web

Author: Bechhofer Sean
Belhajjame Khalid
Corcho Óscar
Garijo Daniel
Goble Carole
Gómez-Pérez José-Manuel
Hettne Kristina
Klyne Graham
Palma Raul
Zhao Jun
Publication venue
Publication date: 03/02/2014
Field of study

Research in life sciences is increasingly being conducted in a digital and online environment. In particular, life scientists have been pioneers in embracing new computational tools to conduct their investigations. To support the sharing of digital objects produced during such research investigations, we have witnessed in the last few years the emergence of specialized repositories, e.g., DataVerse and FigShare. Such repositories provide users with the means to share and publish datasets that were used or generated in research investigations. While these repositories have proven their usefulness, interpreting and reusing evidence for most research results is a challenging task. Additional contextual descriptions are needed to understand how those results were generated and/or the circumstances under which they were concluded. Because of this, scientists are calling for models that go beyond the publication of datasets to systematically capture the life cycle of scientific investigations and provide a single entry point to access the information about the hypothesis investigated, the datasets used, the experiments carried out, the results of the experiments, the people involved in the research, etc. In this paper we present the Research Object (RO) suite of ontologies, which provide a structured container to encapsulate research data and methods along with essential metadata descriptions. Research Objects are portable units that enable the sharing, preservation, interpretation and reuse of research investigation results. The ontologies we present have been designed in the light of requirements that we gathered from life scientists. They have been built upon existing popular vocabularies to facilitate interoperability. Furthermore, we have developed tools to support the creation and sharing of Research Objects, thereby promoting and facilitating their adoption.Comment: 20 page

arXiv.org e-Print Archive

The University of Manchester - Institutional Repository

The Research Object Suite of Ontologies: Sharing and Exchanging Research Data and Methods on the Open Web

Author: Bechhofer Sean
Belhajjame Khalid
Corcho Oscar
Garijo Daniel
Goble Carole
Gomez-Perez Jose Manuel
Hettne Kristina
Klyne Graham
Palma Raul
Zhao Jun
Publication venue
Publication date: 04/02/2014
Field of study

The University of Manchester - Institutional Repository

Evaluating BERT-based scientific relation classifiers for scholarly knowledge graph construction on digital library collections

Author: Auer Sören
Downie J. Stephen
D’Souza Jennifer
Jiang Ming
Publication venue: Heidelberg : Springer
Publication date: 01/01/2021
Field of study

The rapid growth of research publications has placed great demands on digital libraries (DL) for advanced information management technologies. To cater to these demands, techniques relying on knowledge-graph structures are being advocated. In such graph-based pipelines, inferring semantic relations between related scientific concepts is a crucial step. Recently, BERT-based pre-trained models have been popularly explored for automatic relation classification. Despite significant progress, most of them were evaluated in different scenarios, which limits their comparability. Furthermore, existing methods are primarily evaluated on clean texts, which ignores the digitization context of early scholarly publications in terms of machine scanning and optical character recognition (OCR). In such cases, the texts may contain OCR noise, in turn creating uncertainty about existing classifiers’ performances. To address these limitations, we started by creating OCR-noisy texts based on three clean corpora. Given these parallel corpora, we conducted a thorough empirical evaluation of eight Bert-based classification models by focusing on three factors: (1) Bert variants; (2) classification strategies; and, (3) OCR noise impacts. Experiments on clean data show that the domain-specific pre-trained Bert is the best variant to identify scientific relations. The strategy of predicting a single relation each time outperforms the one simultaneously identifying multiple relations in general. The optimal classifier’s performance can decline by around 10% to 20% in F-score on the noisy corpora. Insights discussed in this study can help DL stakeholders select techniques for building optimal knowledge-graph-based systems

Institutionelles Repositorium der Leibniz Universität Hannover

The STEM-ECR Dataset: Grounding Scientific Entity References in STEM Scholarly Content to Authoritative Encyclopedic and Lexicographic Sources

Author: Auer Sören
Brack Arthur
D'Souza Jennifer
Ewerth Ralph
Hoppe Anett
Jaradeh Mohamad Yaser
Publication venue
Publication date: 01/01/2020
Field of study

We introduce the STEM (Science, Technology, Engineering, and Medicine) Dataset for Scientific Entity Extraction, Classification, and Resolution, version 1.0 (STEM-ECR v1.0). The STEM-ECR v1.0 dataset has been developed to provide a benchmark for the evaluation of scientific entity extraction, classification, and resolution tasks in a domain-independent fashion. It comprises abstracts in 10 STEM disciplines that were found to be the most prolific ones on a major publishing platform. We describe the creation of such a multidisciplinary corpus and highlight the obtained findings in terms of the following features: 1) a generic conceptual formalism for scientific entities in a multidisciplinary scientific context; 2) the feasibility of the domain-independent human annotation of scientific entities under such a generic formalism; 3) a performance benchmark obtainable for automatic extraction of multidisciplinary scientific entities using BERT-based neural models; 4) a delineated 3-step entity resolution procedure for human annotation of the scientific entities via encyclopedic entity linking and lexicographic word sense disambiguation; and 5) human evaluations of Babelfy returned encyclopedic links and lexicographic senses for our entities. Our findings cumulatively indicate that human annotation and automatic learning of multidisciplinary scientific concepts as well as their semantic disambiguation in a wide-ranging setting as STEM is reasonable.Comment: Published in LREC 2020. Publication URL https://www.aclweb.org/anthology/2020.lrec-1.268/; Dataset DOI https://doi.org/10.25835/001754

arXiv.org e-Print Archive

Repositorium für Naturwissenschaften und Technik

The STEM-ECR Dataset: Grounding Scientific Entity References in STEM Scholarly Content to Authoritative Encyclopedic and Lexicographic Sources

Author: Auer Sören
Blache Philippe
Brack Arthur
Béchet Frédéric
Calzolari Nicoletta
Choukri Khalid
Cieri Christopher
D'Souza Jennifer
Declerck Thierry
Ewerth Ralph
Goggi Sara
Hoppe Anett
Isahara Hitoshi
Jaradeh Mohamad Yaser
Maegaard Bente
Mariani Joseph
Mazo Hélène
Moreno Asuncion
Odijk Jan
Piperidis Stelios
Publication venue: Paris : The European Language Resources Association (ELRA)
Publication date: 01/01/2020
Field of study

Institutionelles Repositorium der Leibniz Universität Hannover

Connecting TEI Content Into an Ontology of the Editorial Domain

Author: Boot Peter
Koolen Marijn
Publication venue: 'Antibodypedia'
Publication date: 01/01/2021
Field of study

We argued elsewhere that, in order to support interoperable annotations, editions should provide machine-readable identifiers for text and text fragments, as well as information about the text fragments’ type and structure. That is to say, they should be embedded in a Linked Open Data context that facilitates interchange and interpretation of annotation. In this article, assuming a TEI context, we consider the practical question of how the relevant RDF triples are to be derived. How is the edition to know which URIs are to be assigned to which elements in the XML hierarchy, and what are the relevant classes and properties? We discuss different options. Our preference is to generate the relevant triples upon ingestion of the XML file in a version control system and then to store the triples in the TEI xenodata element. We briefly consider situations in which cases the fine-grained annotation that we want to facilitate might be appropriate, or not

Kölner UniversitätsPublikationsServer

Research Articles in Simplified HTML: a Web-first format for HTML-based scholarly articles

Author: Alexander
Atkins Jr
Berjon
Bourne
Brooke
Capadisli
Capadisli
Carlisle
Clark
Constantin
Cyganiak
Di Iorio
Di Iorio
Di Iorio
Di Iorio
Di Mirri
Diggs
Gamma
Gandon
Gao
Garrish
Hickson
Kay
Lin
National Information Standards Organization
Osborne
Peroni
Peroni
Peroni
Peroni
Pettifer
Prud’hommeaux
Raggett
Shotton
Spinaci
Sporny
Sporny
Walsh
Publication venue: 'PeerJ'
Publication date: 10/10/2016
Field of study

Purpose. This paper introduces the Research Articles in Simplified HTML (or RASH), which is a Web-first format for writing HTML-based scholarly papers; it is accompanied by the RASH Framework, a set of tools for interacting with RASH-based articles. The paper also presents an evaluation that involved authors and reviewers of RASH articles submitted to the SAVE-SD 2015 and SAVE-SD 2016 workshops. Design. RASH has been developed aiming to: be easy to learn and use; share scholarly documents (and embedded semantic annotations) through the Web; support its adoption within the existing publishing workflow. Findings. The evaluation study confirmed that RASH is ready to be adopted in workshops, conferences, and journals and can be quickly learnt by researchers who are familiar with HTML. Research Limitations. The evaluation study also highlighted some issues in the adoption of RASH, and in general of HTML formats, especially by less technically savvy users. Moreover, additional tools are needed, e.g., for enabling additional conversions from/to existing formats such as OpenXML. Practical Implications. RASH (and its Framework) is another step towards enabling the definition of formal representations of the meaning of the content of an article, facilitating its automatic discovery, enabling its linking to semantically related articles, providing access to data within the article in actionable form, and allowing integration of data between papers. Social Implications. RASH addresses the intrinsic needs related to the various users of a scholarly article: researchers (focussing on its content), readers (experiencing new ways for browsing it), citizen scientists (reusing available data formally defined within it through semantic annotations), publishers (using the advantages of new technologies as envisioned by the Semantic Publishing movement). Value. RASH helps authors to focus on the organisation of their texts, supports them in the task of semantically enriching the content of articles, and leaves all the issues about validation, visualisation, conversion, and semantic data extraction to the various tools developed within its Framework

Crossref

Directory of Open Access Journals

Open Research Online (The Open University)

Archivio istituzionale della ricerca - Alma Mater Studiorum Università di Bologna

Archivio istituzionale della ricerca - Università di Modena e Reggio Emilia

Hypotheses, evidence and relationships: The HypER approach for representing scientific knowledge claims

Author: Buckingham Shum S.
Carusi A.
de Waard A.
Park J.
Samwald M.
Sándor Á.
Publication venue
Publication date: 01/01/2009
Field of study

Biological knowledge is increasingly represented as a collection of (entity-relationship-entity) triplets. These are queried, mined, appended to papers, and published. However, this representation ignores the argumentation contained within a paper and the relationships between hypotheses, claims and evidence put forth in the article. In this paper, we propose an alternate view of the research article as a network of 'hypotheses and evidence'. Our knowledge representation focuses on scientific discourse as a rhetorical activity, which leads to a different direction in the development of tools and processes for modeling this discourse. We propose to extract knowledge from the article to allow the construction of a system where a specific scientific claim is connected, through trails of meaningful relationships, to experimental evidence. We discuss some current efforts and future plans in this area

CiteSeerX

Open Research Online (The Open University)

Copenhagen University Research Information System