Linguistically annotated multilingual comparable corpora of parliamentary debates ParlaMint.ana 4.0

Agerri, Rodrigo; Agirrezabal, Manex; Agnoloni, Tommaso; Aires, José; Albini, Monica; Alkorta, Jon; Antiba-Cartazo, Iván; Arrieta, Ekain; Barcala, Mario; Bardanca, Daniel; Barkarson, Starkaður; Bartolini, Roberto; Battistoni, Roberto; Bel, Nuria; Bonet Ramos, Maria del Mar; Calzada Pérez, María; Cardoso, Aida; Coole, Matthew; Darģis, Roberts; de Does, Jesse; de Libano, Ruben; Depoorter, Griet; Depuydt, Katrien; Diwersy, Sascha; Dodé, Réka; Erjavec, Tomaž; Fernandez, Kike; Fernández Rei, Elisa; Fišer, Darja; Frontini, Francesca; Garcia, Marcos; García Díaz, Noelia; García Louzao, Pedro; Gavriilidou, Maria; Gkoumas, Dimitris; Grigorov, Ilko; Grigorova, Vladislava; Haltrup Hansen, Dorte; Iruskieta, Mikel; Jarlbrink, Johan; Jelencsik-Mátyus, Kinga; Jongejan, Bart; Kahusk, Neeme; Kirnbauer, Martin; Kopp, Matyáš; Kryvenko, Anna; Ligeti-Nagy, Noémi; Ljubešić, Nikola; Luxardo, Giancarlo; Magariños, Carmen; Magnusson, Måns; Marchetti, Carlo; Marx, Maarten; Meden, Katja; Mendes, Amália; Mochtak, Michal; Montemagni, Simonetta; Mölder, Martin; Navarretta, Costanza; Nitoń, Bartłomiej; Norén, Fredrik Mohammadi; Nwadukwe, Amanda; Ogrodniczuk, Maciej; Ojsteršek, Mihael; Osenova, Petya; Pančur, Andrej; Papavassiliou, Vassilis; Pereira, Rui; Piperidis, Stelios; Pirker, Hannes; Pisani, Marilina; Pol, Henk van der; Prokopidis, Prokopis; Pérez Lago, María; Quochi, Valeria; Rayson, Paul; Regueira, Xosé Luís; Rudolf, Michał; Ruisi, Manuela; Rupnik, Peter; Schopper, Daniel; Simov, Kiril; Sinikallio, Laura; Skubic, Jure; Tamper, Minna; Tungland, Lars Magne; Tuominen, Jouni; van Heusden, Ruben; Varga, Zsófia; Venturi, Giulia; Vidal Miguéns, Adrián; Vider, Kadri; Vivel Couso, Ainhoa; Vladu, Adina Ioana; Vázquez Abuín, Marta; Wissik, Tanja; Yrjänäinen, Väinö; Zevallos, Rodolfo; Çöltekin, Çağrı

Linguistically annotated multilingual comparable corpora of parliamentary debates ParlaMint.ana 4.0

Authors: Rodrigo Agerri
Manex Agirrezabal
Tommaso Agnoloni
José Aires
Monica Albini
Jon Alkorta
Iván Antiba-Cartazo
Ekain Arrieta
Mario Barcala
Daniel Bardanca
Starkaður Barkarson
Roberto Bartolini
Roberto Battistoni
Nuria Bel
Maria del Mar Bonet Ramos
María Calzada Pérez
Aida Cardoso
Matthew Coole
Roberts Darģis
Jesse de Does
Ruben de Libano
Griet Depoorter
Katrien Depuydt
Sascha Diwersy
Réka Dodé
Tomaž Erjavec
Kike Fernandez
Elisa Fernández Rei
Darja Fišer
Francesca Frontini
Marcos Garcia
Noelia García Díaz
Pedro García Louzao
Maria Gavriilidou
Dimitris Gkoumas
Ilko Grigorov
Vladislava Grigorova
Dorte Haltrup Hansen
Mikel Iruskieta
Johan Jarlbrink
Kinga Jelencsik-Mátyus
Bart Jongejan
Neeme Kahusk
Martin Kirnbauer
Matyáš Kopp
Anna Kryvenko
Noémi Ligeti-Nagy
Nikola Ljubešić
Giancarlo Luxardo
Carmen Magariños
Måns Magnusson
Carlo Marchetti
Maarten Marx
Katja Meden
Amália Mendes
Michal Mochtak
Simonetta Montemagni
Martin Mölder
Costanza Navarretta
Bartłomiej Nitoń
Fredrik Mohammadi Norén
Amanda Nwadukwe
Maciej Ogrodniczuk
Mihael Ojsteršek
Petya Osenova
Andrej Pančur
Vassilis Papavassiliou
Rui Pereira
Stelios Piperidis
Hannes Pirker
Marilina Pisani
Henk van der Pol
Prokopis Prokopidis
María Pérez Lago
Valeria Quochi
Paul Rayson
Xosé Luís Regueira
Michał Rudolf
Manuela Ruisi
Peter Rupnik
Daniel Schopper
Kiril Simov
Laura Sinikallio
Jure Skubic
Minna Tamper
Lars Magne Tungland
Jouni Tuominen
Ruben van Heusden
Zsófia Varga
Giulia Venturi
Adrián Vidal Miguéns
Kadri Vider
Ainhoa Vivel Couso
Adina Ioana Vladu
Marta Vázquez Abuín
Tanja Wissik
Väinö Yrjänäinen
Rodolfo Zevallos
Çağrı Çöltekin
Publication date: 24 October 2023
Publisher: CLARIN ERIC

Abstract

ParlaMint 4.0 is a set of comparable corpora containing transcriptions of parliamentary debates of 29 European countries and autonomous regions, mostly starting in 2015 and extending to mid-2022. The individual corpora comprise between 9 and 126 million words and the complete set contains over 1.1 billion words. The transcriptions are divided by days with information on the term, session and meeting, and contain speeches marked by the speaker and their role (e.g. chair, regular speaker). The speeches also contain marked-up transcriber comments, such as gaps in the transcription, interruptions, applause, etc. The corpora have extensive metadata, most importantly on speakers (name, gender, MP and minister status, party affiliation), the political parties and parliamentary groups (name, coalition/opposition status, Wikipedia-sourced left-to-right political orientation, and CHES variables, https://www.chesdata.eu/). Note that some corpora have further metadata, e.g. the year of birth of the speakers, links to their Wikipedia articles, their membership in various committees, etc. The transcriptions are also marked with the subcorpus they belong to ("reference", until 2020-01-30, "covid", from 2020-01-31, and "war", from 2022-02-24). The corpora are encoded according to the Parla-CLARIN TEI recommendation (https://clarin-eric.github.io/parla-clarin/), but have been encoded against the compatible, but much stricter ParlaMint encoding guidelines (https://clarin-eric.github.io/ParlaMint/) and schemas (included in the distribution). The ParlaMint.ana linguistic annotation includes tokenization; sentence segmentation; lemmatisation; Universal Dependencies part-of-speech, morphological features, and syntactic dependencies; and the 4-class CoNLL-2003 named entities. Some corpora also have further linguistic annotations, in particular PoS tagging according a language-specific scheme, with their corpus TEI headers giving further details on the annotation vocabularies and tools used. This entry contains the ParlaMint.ana TEI-encoded linguistically annotated corpora; the derived CoNLL-U files along with TSV metadata of the speeches; and the derived vertical files (with their registry file), suitable for use with CQP-based concordancers, such as CWB, noSketch Engine or KonText. Also included is the 4.0 release of the sample data and scripts available at the GitHub repository of the ParlaMint project at https://github.com/clarin-eric/ParlaMint and the log files produced in the process of building the corpora for this release. The log files show e.g. known errors in the corpora, while more information about known problems is available in the (open) issues at the GitHub repository of the project. This entry contains the linguistically marked-up version of the corpus, while the text version, i.e. without the linguistic annotation is available at http://hdl.handle.net/11356/1859. As opposed to the previous version 3.0, this version adds corpora for Spain (ES), Finland (FI) and the Basque Country (ES-PV); extends the corpora for Austria (AT), Czechia (CZ), Hungary (HU), and Ukraine (UA) with more recent data; adds metadata to political parties and parliamentary groups on left-to-right political orientation from Wikipedia, as well as CHES variables; adds the information on whether a speaker was a minister and when for the corpora that previously lacked this information. The TEI encoding of some details has also changed, and many errors found in 3.0 corpora have been corrected. Furthermore, the vertical files (and hence the individual corpora available on the concordancers) have their meta-data in the local language of the corpus, and not English

Similar works

Full text

Available Versions

Common Language Resources and Technology Infrastructure - Slovenia

oai:www.clarin.si:11356/1860

Last time updated on 27/10/2023