Automatic quality evaluation of Web information is a task with many fields of
applications and of great relevance, especially in critical domains like the
medical one. We move from the intuition that the quality of content of medical
Web documents is affected by features related with the specific domain. First,
the usage of a specific vocabulary (Domain Informativeness); then, the adoption
of specific codes (like those used in the infoboxes of Wikipedia articles) and
the type of document (e.g., historical and technical ones). In this paper, we
propose to leverage specific domain features to improve the results of the
evaluation of Wikipedia medical articles. In particular, we evaluate the
articles adopting an "actionable" model, whose features are related to the
content of the articles, so that the model can also directly suggest strategies
for improving a given article quality. We rely on Natural Language Processing
(NLP) and dictionaries-based techniques in order to extract the bio-medical
concepts in a text. We prove the effectiveness of our approach by classifying
the medical articles of the Wikipedia Medicine Portal, which have been
previously manually labeled by the Wiki Project team. The results of our
experiments confirm that, by considering domain-oriented features, it is
possible to obtain sensible improvements with respect to existing solutions,
mainly for those articles that other approaches have less correctly classified.
Other than being interesting by their own, the results call for further
research in the area of domain specific features suitable for Web data quality
assessment

B Stvilia

DMW Powers

E Marzini

F Cabitza

G Pasi

K Wecel

K Wu

M Hall

NV Chawla

O Bodenreider

SA Azer

TL Saaty

TM Cover

English

arXiv

Automatic quality evaluation of Web information is a task with many fields of applications and of great relevance, especially in critical domains, like the medical one. We move from the intuition that the quality of content of medical Web documents is affected by features related with the specific domain. First, the usage of a specific vocabulary (Domain Informativeness); then, the adoption of specific codes (like those used in the infoboxes of Wikipedia articles) and the type of document (e.g., historical and technical ones). In this paper, we propose to leverage specific domain features to improve the results of the evaluation of Wikipedia medical articles, relying on Natural Language Processing (NLP) and dictionaries-based techniques. The results of our experiments confirm that, by considering domain-oriented features, it is possible to improve existing solutions, mainly with those articles that other approaches have less correctly classified

Cozza, Vittoria

Petrocchi, Marinella

Spognardi, Angelo

Online Research Database In Technology

A Matter of Words - DTU Orbit (09/11/2017) A Matter of Words: NLP for Quality Evaluation of Wikipedia Medical ArticlesAutomatic quality evaluation of Web information is a task with many fields of applications and of great relevance,especially in critical domains, like the medical one. We move from the intuition that the quality of content of medical Webdocuments is affected by features related with the specific domain. First, the usage of a specific vocabulary (DomainInformativeness); then, the adoption of specific codes (like those used in the infoboxes of Wikipedia articles) and the typeof document (e.g., historical and technical ones). In this paper, we propose to leverage specific domain features toimprove the results of the evaluation of Wikipedia medical articles, relying on Natural Language Processing (NLP) anddictionaries-based techniques. The results of our experiments confirm that, by considering domain-oriented features, it ispossible to improve existing solutions, mainly with those articles that other approaches have less correctly classified. General informationState: PublishedOrganisations: Department of Applied Mathematics and Computer Science , Embedded Systems Engineering, IIT-CNRAuthors: Cozza, V. (Ekstern), Petrocchi, M. (Ekstern), Spognardi, A. (Intern)Number of pages: 9Pages: 448-456Publication date: 2016 Host publication informationTitle of host publication: Web Engineering : 16th International Conference, ICWE 2016, Lugano, Switzerland, June 6-9,2016. ProceedingsVolume: 9671Publisher: SpringerISBN (Print): 978-3-319-38790-1ISBN (Electronic): 978-3-319-38791-8 Series: Lecture Notes in Computer ScienceISSN: 0302-9743Main Research Area: Technical/natural sciencesConference: The 16th International Conference on Web Engineering (ICWE2016), USI Lugano, Switzerland, 06/06/2016 -06/06/2016Information Systems Applications (incl. Internet), Information Storage and Retrieval, Software Engineering, ComputerAppl. in Administrative Data Processing, User Interfaces and Human Computer Interaction, Artificial Intelligence (incl.Robotics)DOIs: 10.1007/978-3-319-38791-8_31 Source: FindItSource-ID: 2306622913Publication: Research - peer-review › Article in proceedings – Annual report year: 2016 

A matter of words: NLP for quality evaluation of Wikipedia medical articles

Abstract

Similar works

Full text

Available Versions

Online Research Database In Technology

Archivio della ricerca- Università di Roma La Sapienza

Catalogo dei prodotti della ricerca

Archivio istituzionale della ricerca - Università di Padova

Crossref