Crowdsourcing Image Extraction and Annotation: Software Development and Case Study

Bennett, Carl; Berardi, Vincent; Brennan, Kathleen P.J.; Cornejo, Aisha; Harlan, John; Jofre, Ana

Crowdsourcing Image Extraction and Annotation: Software Development and Case Study

Authors: Carl Bennett
Vincent Berardi
Kathleen P.J. Brennan
Aisha Cornejo
John Harlan
Ana Jofre
Publication date: 1 January 2020
Publisher: Chapman University Digital Commons

Abstract

We describe the development of web-based software that facilitates large-scale, crowdsourced image extraction and annotation within image-heavy corpora that are of interest to the digital humanities. An application of this software is then detailed and evaluated through a case study where it was deployed within Amazon Mechanical Turk to extract and annotate faces from the archives of Time magazine. Annotation labels included categories such as age, gender, and race that were subsequently used to train machine learning models. The systemization of our crowdsourced data collection and worker quality verification procedures are detailed within this case study. We outline a data verification methodology that used validation images and required only two annotations per image to produce high-fidelity data that has comparable results to methods using five annotations per image. Finally, we provide instructions for customizing our software to meet the needs for other studies, with the goal of offering this resource to researchers undertaking the analysis of objects within other image-heavy archives

Similar works

Full text

Open in the Core reader

Download PDF

Available Versions

Chapman University Digital Commons

oai:digitalcommons.chapman.edu...

Last time updated on 02/03/2021