Automated document metadata extraction

Bolanle Adefowoke Ojokoh; Olumide Sunday Adewale; Samuel Oluwole Falaki

text

oai:CiteSeerX.psu:10.1.1.955.4926

Automated document metadata extraction

Authors: Bolanle Adefowoke Ojokoh
Olumide Sunday Adewale
Samuel Oluwole Falaki
Publication date: 1 January 2009
Publisher
Doi

Abstract

Web documents are available in various forms, most of which do not carry additional semantics. This paper presents a model for general document metadata extraction. The model, which combines segmentation by keywords and pattern matching techniques, was implemented using PHP, MySQL, JavaScript and HTML. The system was tested with 40 randomly selected PDF documents (mainly theses). An evaluation of the sys-tem was done using standard criteria measures namely precision, recall, accuracy and F-measure. The results show that the model is relatively effective for the task of metadata extraction, especially for theses and dissertations. A combination of machine learning with these rule-based methods will be explored in the future for better results

Similar works

Full text

CiteSeerX

oai:CiteSeerX.psu:10.1.1.955.4...

Last time updated on 01/11/2017

This paper was published in CiteSeerX.

Having an issue?

Is data on this page outdated, violates copyrights or anything else? Report the problem now and we will take corresponding actions after reviewing your request.