Search CORE

2,299 research outputs found

Data formats for phonological corpora

Author: Romary Laurent
Witt Andreas
Publication venue
Publication date: 01/01/2012
Field of study

The goal of the present chapter is to explore the possibility of providing the research (but also the industrial) community that commonly uses spoken corpora with a stable portfolio of well-documented standardised formats that allow a high re-use rate of annotated spoken resources and, as a consequence, better interoperability across tools used to produce or exploit such resources.Comment: Handbook of Corpus Phonology Oxford University Press (Ed.) (2012

arXiv.org e-Print Archive

INRIA a CCSD electronic archive server

Enroller: an experiment in aggregating resources

Author: Anderson J.
Publication venue: Editions Rodopi B.V.
Publication date: 01/01/2013
Field of study

This chapter describes a collaborative project between e-scientists and humanists working to create an online repository of linguistic data sets and tools. Corpora, dictionaries, and a thesaurus are brought together to enable a new method of research. It combines our most advanced knowledge in both computing and linguistic research techniques

Enlighten

A new web interface to facilitate access to corpora: development of the ASLLRP data access interface

Author: Neidle Carol
Vogler Christian
Publication venue
Publication date: 01/01/2012
Field of study

A significant obstacle to broad utilization of corpora is the difficulty in gaining access to the specific subsets of data and annotations that may be relevant for particular types of research. With that in mind, we have developed a web-based Data Access Interface (DAI), to provide access to the expanding datasets of the American Sign Language Linguistic Research Project (ASLLRP). The DAI facilitates browsing the corpora, viewing videos and annotations, searching for phenomena of interest, and downloading selected materials from the website. The web interface, compared to providing videos and annotation files off-line, also greatly increases access by people that have no prior experience in working with linguistic annotation tools, and it opens the door to integrating the data with third-party applications on the desktop and in the mobile space. In this paper we give an overview of the available videos, annotations, and search functionality of the DAI, as well as plans for future enhancements. We also summarize best practices and key lessons learned that are crucial to the success of similar projects

CiteSeerX

Boston University Institutional Repository (OpenBU)

The Validation of Speech Corpora

Author: Baumann Angela
Draxler Christoph
Ellbogen Tania
Hoole Phil
Schiel Florian
Steffen Alexander
Publication venue
Publication date: 01/01/2012
Field of study

1.2 Intended audience........................

CiteSeerX

Open Access LMU

A Formal Framework for Linguistic Annotation

Author: Bird Steven
Liberman Mark
Publication venue
Publication date: 01/01/1999
Field of study

`Linguistic annotation' covers any descriptive or analytic notations applied to raw language data. The basic data may be in the form of time functions -- audio, video and/or physiological recordings -- or it may be textual. The added notations may include transcriptions of all sorts (from phonetic features to discourse structures), part-of-speech and sense tagging, syntactic analysis, `named entity' identification, co-reference annotation, and so on. While there are several ongoing efforts to provide formats and tools for such annotations and to publish annotated linguistic databases, the lack of widely accepted standards is becoming a critical problem. Proposed standards, to the extent they exist, have focussed on file formats. This paper focuses instead on the logical structure of linguistic annotations. We survey a wide variety of existing annotation formats and demonstrate a common conceptual core, the annotation graph. This provides a formal framework for constructing, maintaining and searching linguistic annotations, while remaining consistent with many alternative data structures and file formats.Comment: 49 page

arXiv.org e-Print Archive

CiteSeerX

ScholarlyCommons@Penn

Annotation Graphs and Servers and Multi-Modal Resources: Infrastructure for Interdisciplinary Education, Research and Development

Author: Bird Steven
Cieri Christopher
Publication venue
Publication date: 01/01/2001
Field of study

Annotation graphs and annotation servers offer infrastructure to support the analysis of human language resources in the form of time-series data such as text, audio and video. This paper outlines areas of common need among empirical linguists and computational linguists. After reviewing examples of data and tools used or under development for each of several areas, it proposes a common framework for future tool development, data annotation and resource sharing based upon annotation graphs and servers.Comment: 8 pages, 6 figure

arXiv.org e-Print Archive

CiteSeerX

The Production of Speech Corpora

Author: Baumann Angela
Draxler Christoph
Ellbogen Tania
Schiel Florian
Steffen Alexander
Publication venue
Publication date: 21/03/2012
Field of study

Open Access LMU

Phonology and intonation

Author: Féry Caroline
Hellmuth Sam
Kügler Frank
Mayer Jörg
Stoel Ruben
Vijver Ruben van de
Publication venue
Publication date: 06/11/2008
Field of study

The encoding standards for phonology and intonation are designed to facilitate consistent annotation of the phonological and intonational aspects of information structure, in languages across a range ofprosodic types. The guidelines are designed with the aim that a nonspecialist in phonology can both implement and interpret the resulting annotation

Hochschulschriftenserver - Universität Frankfurt am Main

Production Methods

Author: Eisenbeiss Sonja
Publication venue: 'John Benjamins Publishing Company'
Publication date: 01/01/2010
Field of study

University of Essex Research Repository

EXMARaLDA - Creating, Analysing and Sharing Spoken Language Corpora for Pragmatic Research

Author: Schmidt Thomas
Wörner Kai
Publication venue
Publication date: 07/05/2014
Field of study

This paper presents EXMARaLDA, a system for the computer-assisted creation and analysis of spoken language corpora. The first part contains some general observations about technological and methodological requirements for doing corpus-based pragmatics. The second part explains the systems architecture and gives an overview of its most important software components a transcription editor, a corpus management tool and a corpus query tool. The last part presents some corpora which have been or are currently being compiled with the help of EXMARaLDA

Publikationsserver des Instituts für Deutsche Sprache