A Diachronic Italian Corpus based on “L’Unità”

Basile, Pierpaolo; Caputo, Annalina; Caselli, Tommaso; Cassotti, Pierluigi; Varvara, Rossella

A Diachronic Italian Corpus based on “L’Unità”

Authors: Pierpaolo Basile
Annalina Caputo
Tommaso Caselli
Pierluigi Cassotti
Rossella Varvara
Publication date: 1 January 2020
Publisher: CEUR Workshop Proceedings (CEUR-WS.org)
Doi

Abstract

In this paper, we describe the creation of a diachronic corpus for Italian by exploiting the digital archive of the newspaper “L’Unità”. We automatically clean and annotate the corpus with PoS tags, lemmas, named entities and syntactic dependencies. Moreover, we compute frequency-based time series for tokens,lemmas and entities. We show some interesting corpus statistics taking into account the temporal dimension and describe some examples of usage of time series