The Shortcomings of Language Tags for Linked Data When Modeling Lesser-Known Languages

Gillis-Webber, Frances; Tittel, Sabine

research

The Shortcomings of Language Tags for Linked Data When Modeling Lesser-Known Languages

Authors: Frances Gillis-Webber
Sabine Tittel
Publication date: 1 January 2019
Publisher: OASIcs - OpenAccess Series in Informatics. 2nd Conference on Language, Data and Knowledge (LDK 2019)
Doi

Abstract

In recent years, the modeling of data from linguistic resources with Resource Description Framework (RDF), following the Linked Data paradigm and using the OntoLex-Lemon vocabulary, has become a prevalent method to create datasets for a multilingual web of data. An important aspect of data modeling is the use of language tags to mark lexicons, lexemes, word senses, etc. of a linguistic dataset. However, attempts to model data from lesser-known languages show significant shortcomings with the authoritative list of language codes by ISO 639: for many lesser-known languages spoken by minorities and also for historical stages of languages, language codes, the basis of language tags, are simply not available. This paper discusses these shortcomings based on the examples of three such languages, i.e., two varieties of click languages of Southern Africa together with Old French, and suggests solutions for the issues identified

Similar works

Full text

Open in the Core reader

Download PDF

Available Versions

DROPS Dagstuhl Research Online Publication Server

oai:drops-oai.dagstuhl.de:1036...

Last time updated on 22/05/2019