Search CORE

3 research outputs found

MultiLexNorm: A Shared Task on Multilingual Lexical Normalization

Author: Baldwin T
Caselli T
Ljubešić N
Mahendra R
Muller B
Plank B
Ramponi A
Roncal ISV
Sidorenko W
van der Goot R
Workshop on Noisy User-Generated Text
Zubiaga A
Çetinoğlu Ö
Çolakoğlu T
Publication venue
Publication date: 01/01/2021
Field of study

Lexical normalization is the task of transforming an utterance into its standardized form. This task is beneficial for downstream analysis, as it provides a way to harmonize (often spontaneous) linguistic variation. Such variation is typical for social media on which information is shared in a multitude of ways, including diverse languages and code-switching. Since the seminal work of Han and Baldwin (2011) a decade ago, lexical normalization has attracted attention in English and multiple other languages. However, there exists a lack of a common benchmark for comparison of systems across languages with a homogeneous data and evaluation setup. The MULTILEXNORM shared task sets out to fill this gap. We provide the largest publicly available multilingual lexical normalization benchmark including 12 language variants. We propose a homogenized evaluation setup with both intrinsic and extrinsic evaluation. As extrinsic evaluation, we use dependency parsing and part-of-speech tagging with adapted evaluation metrics (a-LAS, a-UAS, and a-POS) to account for alignment discrepancies. The shared task hosted at W-NUT 2021 attracted 9 participants and 18 submissions. The results show that neural normalization systems outperform the previous state-of-the-art system by a large margin. Downstream parsing and part-of-speech tagging performance is positively affected but to varying degrees, with improvements of up to 1.72 a-LAS, 0.85 a-UAS, and 1.54 a-POS for the winning system

Queen Mary Research Online

Results of the WNUT16 Named Entity Recognition Shared Task

Author: Alan Ritter
Benjamin Strauss
Bethany Toma
de Marneffe Marie-Catherine
Proceedings of the 2nd Workshop on Noisy User-generated Text (WNUT)
Wei Xu
Publication venue: The COLING 2016 Organizing Committee
Publication date: 01/01/2016
Field of study

This paper presents the results of the Twitter Named Entity Recognition shared task associated with W-NUT 2016: a named entity tagging task with 10 teams participating. We outline the shared task, annotation process and dataset statistics, and provide a high-level overview of the participating systems for each shared task

DIAL UCLouvain

Shared Tasks of the 2015 Workshop on Noisy User-generated Text: Twitter Lexical Normalization and Named Entity Recognition

Author: Baldwin Timothy
de Marneffe Marie-Catherine
Han Bo
Kim Young-Bum
Proceedings of the Workshop on Noisy User-generated Text
Ritter Alan
Xu Wei
Publication venue
Publication date: 01/01/2015
Field of study

DIAL UCLouvain