Shami: A Corpus of Levantine Arabic Dialects

Abu Kwaik, Kathrein; Chatzikyriakidis, Stergios; Dobnik, Simon; Saad, Motaz K

Shami: A Corpus of Levantine Arabic Dialects

Authors: Kathrein Abu Kwaik
Stergios Chatzikyriakidis
Simon Dobnik
Motaz K Saad
Publication date: 1 January 2018
Publisher

Abstract

Modern Standard Arabic (MSA) is the official language used in education and media across the Arab world both in writing and formal speech. However, in daily communication several dialects depending on the country, region as well as other social factors, are used. With the emergence of social media, the dialectal amount of data on the Internet have increased and the NLP tools that support MSA are not well-suited to process this data due to the difference between the dialects and MSA. In this paper, we construct the Shami corpus, the first Levantine Dialect Corpus (SDC) covering data from the four dialects spoken in Palestine, Jordan, Lebanon and Syria. We also describe rules for pre-processing without affecting the meaning so that it is processable by NLP tools. We choose Dialect Identification as the task to evaluate SDC and compare it with two other corpora. In this respect, experiments are conducted using different parameters based on n-gram models and Naive Bayes classifiers. SDC is larger than the existing corpora in terms of size, words and vocabularies. In addition, we use the performance on the Language Identification task to exemplify the similarities and differences in the individual dialects

Similar works

Full text

Available Versions

Institutional Repository of the Islamic University of Gaza

oai:iugspace.iugaza.edu.ps:20....

Last time updated on 19/02/2021