Search CORE

824 research outputs found

Modeling Language Variation and Universals: A Survey on Typological Linguistics for Natural Language Processing

Author: Berzak Yevgeni
Korhonen Anna
O'Horan Helen
Poibeau Thierry
Ponti Edoardo Maria
Reichart Roi
Shutova Ekaterina
Vulić Ivan
Publication venue
Publication date: 27/02/2019
Field of study

Linguistic typology aims to capture structural and semantic variation across the world's languages. A large-scale typology could provide excellent guidance for multilingual Natural Language Processing (NLP), particularly for languages that suffer from the lack of human labeled resources. We present an extensive literature survey on the use of typological information in the development of NLP techniques. Our survey demonstrates that to date, the use of information in existing typological databases has resulted in consistent but modest improvements in system performance. We show that this is due to both intrinsic limitations of databases (in terms of coverage and feature granularity) and under-employment of the typological features included in them. We advocate for a new approach that adapts the broad and discrete nature of typological categories to the contextual and continuous nature of machine learning algorithms used in contemporary NLP. In particular, we suggest that such approach could be facilitated by recent developments in data-driven induction of typological knowledge

arXiv.org e-Print Archive

Edinburgh Research Explorer

Apollo (Cambridge)

UvA-DARE

International Migration, Integration and Social Cohesion online publications

Max-Planck-Institute for Psycholinguistics: Annual Report 2003

Author: Johnson E.
Matsuo A.
Publication venue: MPI for Psycholinguistics
Publication date: 01/01/2003
Field of study

MPG.PuRe

Modeling Language Variation and Universals: A Survey on Typological Linguistics for Natural Language Processing

Author: Berzak Y.
Korhonen A.
O'Horan H.
Poibeau T.
Ponti E.M.
Reichart R.
Shutova E.
Vulić I.
Publication venue: 'MIT Press - Journals'
Publication date: 01/09/2019
Field of study

International Migration, Integration and Social Cohesion online publications

Modeling Language Variation and Universals: A Survey on Typological Linguistics for Natural Language Processing

Author: Berzak Yevgeni
Korhonen Anna
O'Horan Helen
Poibeau Thierry
Ponti Edoardo Maria
Reichart Roi
Shutova Ekaterina
Vulic Ivan
Publication venue: COMPUTATIONAL LINGUISTICS
Publication date: 09/08/2018
Field of study

Linguistic typology aims to capture structural and semantic variation across the world’s languages. A large-scale typology could provide excellent guidance for multilingual Natural Language Processing (NLP), particularly for languages that suffer from the lack of human labeled resources. We present an extensive literature survey on the use of typological information in the development of NLP techniques. Our survey demonstrates that to date, the use of information in existing typological databases has resulted in consistent but modest improvements in system performance. We show that this is due to both intrinsic limitations of databases (in terms of coverage and feature granularity) and under-utilization of the typological features included in them. We advocate for a new approach that adapts the broad and discrete nature of typological categories to the contextual and continuous nature of machine learning algorithms used in contemporary NLP. In particular, we suggest that such an approach could be facilitated by recent developments in data-driven induction of typological knowledge.</jats:p

arXiv.org e-Print Archive

Edinburgh Research Explorer

Apollo (Cambridge)

UvA-DARE

International Migration, Integration and Social Cohesion online publications

Proceedings of the Seventh Italian Conference on Computational Linguistics CLiC-it 2020

Author
Publication venue: 'OpenEdition'
Publication date: 01/07/2022
Field of study

On behalf of the Program Committee, a very warm welcome to the Seventh Italian Conference on Computational Linguistics (CLiC-it 2020). This edition of the conference is held in Bologna and organised by the University of Bologna. The CLiC-it conference series is an initiative of the Italian Association for Computational Linguistics (AILC) which, after six years of activity, has clearly established itself as the premier national forum for research and development in the fields of Computational Linguistics and Natural Language Processing, where leading researchers and practitioners from academia and industry meet to share their research results, experiences, and challenges

Directory of Open Access Books (DOAB)

The evolution of language: Proceedings of the Joint Conference on Language Evolution (JCoLE)

Author
Publication venue: Joint Conference on Language Evolution (JCoLE)
Publication date: 01/01/2022
Field of study

MPG.PuRe

FOUND IN SPACE: A CROSS-LINGUISTIC ANALYSIS OF SECOND LANGUAGE LEARNERS IN ENGLISH MAP TASK PERFORMANCE

Author: Metheny Susan K
Publication venue: UNM Digital Repository
Publication date: 01/07/2019
Field of study

Understanding the relationship between first and second language use in the area of spatial language has broader implications for our understanding of language learning and consequences for the construction of bilingual assessment instruments for second language learners. This study shows that observing and interpreting the task of map drawing and the related behavior of explaining maps can be a way to explore the linguistic emergence of the conceptualization of spatial language (at a moment of simultaneous and synchronized incarnation). Altogether, 50 dyads (pairs) participated in the New Mexico Map Task Project; the project included native speakers of English, Russian, Japanese, Navajo, and Spanish. In an examination of how the grammatical constructions used for spatial descriptions in a speaker\u27s first language carry over into the usage of this speaker\u27s second language, new observations include the intra-subject comparison of dyadic map task performances. Each non-native English-speaking dyad participates in two map task performances: one in their native language and one in their second language, English. Evidence was generated through morphosyntactic, phonological, and pragmatic analyses performed on the sound files of the transcripts. This evidence confirms the connection between the participants\u27 productions of tokens of selected landmark names both in their native language and their second language. Combining the results of linguistic analyses with educational assessment frameworks predicts the development of an instrument for use with immigrant and refugee students from areas of conflict

Recommended from our members

Sociolinguistically Driven Approaches for Just Natural Language Processing

Author: Blodgett Su Lin
Publication venue: ScholarWorks@UMass Amherst
Publication date: 06/04/2021
Field of study

Natural language processing (NLP) systems are now ubiquitous. Yet the benefits of these language technologies do not accrue evenly to all users, and indeed they can be harmful; NLP systems reproduce stereotypes, prevent speakers of non-standard language varieties from participating fully in public discourse, and re-inscribe historical patterns of linguistic stigmatization and discrimination. How harms arise in NLP systems, and who is harmed by them, can only be understood at the intersection of work on NLP, fairness and justice in machine learning, and the relationships between language and social justice. In this thesis, we propose to address two questions at this intersection: i) How can we conceptualize harms arising from NLP systems?, and ii) How can we quantify such harms? We propose the following contributions. First, we contribute a model in order to collect the first large dataset of African American Language (AAL)-like social media text. We use the dataset to quantify the performance of two types of NLP systems, identifying disparities in model performance between Mainstream U.S. English (MUSE)- and AAL-like text. Turning to the landscape of bias in NLP more broadly, we then provide a critical survey of the emerging literature on bias in NLP and identify its limitations. Drawing on work across sociology, sociolinguistics, linguistic anthropology, social psychology, and education, we provide an account of the relationships between language and injustice, propose a taxonomy of harms arising from NLP systems grounded in those relationships, and propose a set of guiding research questions for work on bias in NLP. Finally, we adapt the measurement modeling framework from the quantitative social sciences to effectively evaluate approaches for quantifying bias in NLP systems. We conclude with a discussion of recent work on bias through the lens of style in NLP, raising a set of normative questions for future work

ScholarWorks@UMass Amherst

Application of self-organizing maps to multilingual text mining (Arabic-English).

Author: Saleh Abdulsamad A. M.
Publication venue: 'De Montfort University'
Publication date: 01/01/2008
Field of study

De Montfort University Open Research Archive