Search CORE

628 research outputs found

A CHAT-Based Annotation Scheme for Case and Noun-Phrase Inflection in Child Language Data

Author: Eisenbeiss Sonja
Sonnenstuhl Ingrid
Publication venue: Essex Research Reports in Linguistics
Publication date: 01/01/2011
Field of study

University of Essex Research Repository

Kölner UniversitätsPublikationsServer

Proceedings of the workshop on language technology for normalisation of less-resourced languages (SaLTMiL 8 - AfLaT 2012)

Author: De Pauw Guy
de Schryver Gilles-Maurice
Forcada Mike L
Sarasola Kepa
Tyers Francis M
Wagacha Peter W
Publication venue: European Language Resources Association
Publication date: 01/01/2012
Field of study

Ghent University Academic Bibliography

LatMor: A Latin Finite-State Morphology Encoding Vowel Quantity

Author: Najock Dietmar
Schmid Helmut
Springmann Uwe
Publication venue: 'Walter de Gruyter GmbH'
Publication date: 01/01/2016
Field of study

We present the first large-coverage finite-state open-source morphology for Latin (called LatMor) which parses as well as generates vowel quantity information. LatMor is based on the Berlin Latin Lexicon comprising about 70,000 lemmata of classical Latin compiled by the group of Dietmar Najock in their work on concordances of Latin authors (see Rapsch and Najock, 1991) which was recently updated by us. Compared to the well-known Morpheus system of Crane (1991, 1998), which is written in the C programming language, based on 50,000 lemmata of Lewis and Short (1907), not well documented and therefore not easily extended, our new morphology has a larger vocabulary, is about 60 to 1200 times faster and is built in the form of finite-state transducers which can analyze as well as generate wordforms and represent the state-of-the-art implementation method in computational morphology. The current coverage of LatMor is evaluated against Morpheus and other existing systems (some of which are not openly accessible), and is shown to rank first among all systems together with the Pisa LEMLAT morphology (not yet openly accessible). Recall has been analyzed taking the Latin Dependency Treebank(1) as gold data and the remaining defect classes have been identified. LatMor is available under an open source licence to allow its wide usage by all interested parties

Crossref

Directory of Open Access Journals

Open Access LMU

A very unpredictable ‘person’: A corpus-based approach to suppletion in West Polesian

Author: Roncero K.
Publication venue: 'Peoples'' Friendship University of Russia'
Publication date: 01/01/2022
Field of study

MPG.PuRe

Compiling and annotating a learner corpus for a morphologically rich language: CzeSL, a corpus of non-native Czech

Author: Hana Jiří
Jelínek Tomáš
Rosen Alexandr
Vidová Hladká Barbora
Škodová Svatava
Štindlová Barbora
Publication venue: 'Charles University in Prague, Karolinum Press'
Publication date: 01/01/2020
Field of study

Learner corpora, linguistic collections documenting a language as used by learners, provide an important empirical foundation for language acquisition research and teaching practice. This book presents CzeSL, a corpus of non-native Czech, against the background of theoretical and practical issues in the current learner corpus research. Languages with rich morphology and relatively free word order, including Czech, are particularly challenging for the analysis of learner language. The authors address both the complexity of learner error annotation, describing three complementary annotation schemes, and the complexity of description of non-native Czech in terms of standard linguistic categories. The book discusses in detail practical aspects of the corpus creation: the process of collection and annotation itself, the supporting tools, the resulting data, their formats and search platforms. The chapter on use cases exemplifies the usefulness of learner corpora for teaching, language acquisition research, and computational linguistics. Any researcher developing learner corpora will surely appreciate the concluding chapter listing lessons learned and pitfalls to avoid

CU Digital Repository

Directory of Open Access Books (DOAB)

SenZi: A Sentiment Analysis Lexicon for the Latinised Arabic (Arabizi)

Author: Alani Harith
Fernandez Miriam
Glavas Goran
Hajj Hazem
Sharafeddine Sanaa
Tobaili Taha
Publication venue
Publication date: 01/01/2019
Field of study

Arabizi is an informal written form of dialectal Arabic transcribed in Latin alphanumeric characters. It has a proven popularity on chat platforms and social media, yet it suffers from a severe lack of natural language processing (NLP) resources. As such, texts written in Arabizi are often disregarded in sentiment analysis tasks for Arabic. In this paper we describe the creation of a sentiment lexicon for Arabizi that was enriched with word embeddings. The result is a new Arabizi lexicon consisting of 11.3K positive and 13.3K negative words. We evaluated this lexicon by classifying the sentiment of Arabizi tweets achieving an F1-score of 0.72. We provide a detailed error analysis to present the challenges that impact the sentiment analysis of Arabizi

Crossref

Open Research Online

MAnnheim DOCument Server

Compte rendu

Author: McPherson Laura
Publication venue: 'OpenEdition'
Publication date: 21/12/2017
Field of study

1. Introduction Solomiac’s book is a welcome and much needed addition to the literature on Northwestern Mande languages and languages of the Samogo group in particular. This group, consisting of around six small languages spoken in Mali and Burkina Faso (Dzùùngoo, Duungooma, Bankagooma, Jowulu, Seenku, and Kpeengo), has until now received only scant attention in the literature. Published sources include only a description of the phonology of Jowulu published in the same series (Djilla, Eenkho..

OpenEdition

Incorporating Grammatical Features in the Modeling of the Slovak Language for Continuous Speech Recognition

Author: Daniel Hládek
Ján Staš
Jozef Juhár
Publication venue: 'IntechOpen'
Publication date: 28/11/2012
Field of study

IntechOpen

Recommended from our members

Sentiment Analysis for the Low-Resourced Latinised Arabic "Arabizi"

Author: Tobaili Taha
Publication venue
Publication date: 02/11/2020
Field of study

The expansion of digital communication mediums from private mobile messaging into the public through social media presented an opportunity for the data science research and industry to mine the generated big data for artificial information extraction. A popular information extraction task is sentiment analysis, which aims at extracting polarity opinions, positive, negative, or neutral, from the written natural language. This science helped organisations better understand the public’s opinion towards events, news, public figures, and products. However, sentiment analysis has advanced for the English language ahead of Arabic. While sentiment analysis for Arabic is developing in the literature of Natural Language Processing (NLP), a popular variety of Arabic, Arabizi, has been overlooked for sentiment analysis advancements. Arabizi is an informal transcription of the spoken dialectal Arabic in Latin script used for social texting. It is known to be common among the Arab youth, yet it is overlooked in efforts on Arabic sentiment analysis for its linguistic complexities. As to Arabic, Arabizi is rich in inflectional morphology, but also codeswitched with English or French, and distinctively transcribed without adhering to a standard orthography. The rich morphology, inconsistent orthography, and codeswitching challenges are compounded together to have a multiplied effect on the lexical sparsity of the language, where each Arabizi word becomes eligible to be spelled in many ways, that, in addition to the mixing of other languages within the same textual context. The resulting high degree of lexical sparsity defies the very basics of sentiment analysis, classification of positive and negative words. Arabizi is even faced with a severe shortage of data resources that are required to set out any sentiment analysis approach. In this thesis, we tackle this gap by conducting research on sentiment analysis for Arabizi. We addressed the sparsity challenge by harvesting Arabizi data from multi-lingual social media text using deep learning to build Arabizi resources for sentiment analysis. We developed six new morphologically and orthographically rich Arabizi sentiment lexicons and set the baseline for Arabizi sentiment analysis on social media

Open Research Online