Search CORE

2,225 research outputs found

A Survey of Word Reordering in Statistical Machine Translation: Computational Models and Language Phenomena

Author: Bisazza Arianna
Federico Marcello
Publication venue: 'MIT Press - Journals'
Publication date: 14/03/2016
Field of study

Word reordering is one of the most difficult aspects of statistical machine translation (SMT), and an important factor of its quality and efficiency. Despite the vast amount of research published to date, the interest of the community in this problem has not decreased, and no single method appears to be strongly dominant across language pairs. Instead, the choice of the optimal approach for a new translation task still seems to be mostly driven by empirical trials. To orientate the reader in this vast and complex research area, we present a comprehensive survey of word reordering viewed as a statistical modeling challenge and as a natural language phenomenon. The survey describes in detail how word reordering is modeled within different string-based and tree-based SMT frameworks and as a stand-alone task, including systematic overviews of the literature in advanced reordering modeling. We then question why some approaches are more successful than others in different language pairs. We argue that, besides measuring the amount of reordering, it is important to understand which kinds of reordering occur in a given language pair. To this end, we conduct a qualitative analysis of word reordering phenomena in a diverse sample of language pairs, based on a large collection of linguistic knowledge. Empirical results in the SMT literature are shown to support the hypothesis that a few linguistic facts can be very useful to anticipate the reordering characteristics of a language pair and to select the SMT framework that best suits them.Comment: 44 pages, to appear in Computational Linguistic

arXiv.org e-Print Archive

Crossref

Archivio della ricerca - Fondazione Bruno Kessler

UvA-DARE

International Migration, Integration and Social Cohesion online publications

The Edinburgh/LMU Hierarchical Machine Translation System for WMT 2016

Author: Fraser Alexander
Haddow Barry
Huck Matthias
Publication venue: 'Association for Computational Linguistics (ACL)'
Publication date: 12/08/2016
Field of study

Edinburgh Research Explorer

Comparison of Data Selection Techniques for the Translation of Video Lectures

Author: Dymetman Marc
Giménez Pastor Adrián
Juan Císcar Alfonso
Martínez-Villaronga Adrià
Mirkin Shashar
Ney Hermann
Servan Christophe
Wuebker Joern
Publication venue: Association for Machine Translation in the Americas
Publication date: 01/01/2014
Field of study

[EN] For the task of online translation of scientific video lectures, using huge models is not possible. In order to get smaller and efficient models, we perform data selection. In this paper, we perform a qualitative and quantitative comparison of several data selection techniques, based on cross-entropy and infrequent n-gram criteria. In terms of BLEU, a combination of translation and language model cross-entropy achieves the most stable results. As another important criterion for measuring translation quality in our application, we identify the number of out-ofvocabulary words. Here, infrequent n-gram recovery shows superior performance. Finally, we combine the two selection techniques in order to benefit from both their strengths.The research leading to these results has received funding from the European Union Seventh Framework Programme (FP7/2007-2013) under grant agreement no 287755 (transLectures), and the Spanish MINECO Active2Trans (TIN2012-31723) research project.Wuebker, J.; Ney, H.; Martínez-Villaronga, A.; Giménez Pastor, A.; Juan Císcar, A.; Servan, C.; Dymetman, M.... (2014). Comparison of Data Selection Techniques for the Translation of Video Lectures. Association for Machine Translation in the Americas. http://hdl.handle.net/10251/54431

RiuNet

Publikationsserver der RWTH Aachen University

Comparison of Data Selection Techniques for the Translation of Video Lectures

Author: Wuebker Joern
Ney Hermann
Martínez-Villaronga Adrià
Giménez Pastor Adrián
Juan Císcar Alfonso
Servan Christophe
Dymetman Marc
Mirkin Shashar
Publication venue: Association for Machine Translation in the Americas
Publication date: 22/10/2014
Field of study

RiuNet

Institutional Repositories DataBase (IRDB)

The QT21/HimL Combined Machine Translation System

Author: Alkhouli Tamer
Allauzen Alexandre
Aufrant Lauriane
Blain Frédéric
Bojar Ondrej
Braune Fabienne
Burlot Franck
Frank Stella
Fraser Alexander
Haddow Barry
Huck Matthias
knyazeva elena
Lavergne Thomas
Ney Hermann
Niehues Jan
Peter Jan-Thorsten
Pinnis Marcis
Sennrich Rico
Specia Lucia
Tamchyna Ales
Waibel Alex
Yvon François
Publication venue: 'Association for Computational Linguistics (ACL)'
Publication date: 01/01/2016
Field of study

This paper describes the joint submission of the QT21 and HimL projects for the English→Romanian translation task of the ACL 2016 First Conference on Machine Translation (WMT 2016). The submission is a system combination which combines twelve different statistical machine translation systems provided by the different groups (RWTH Aachen University, LMU Munich, Charles University in Prague, University of Edinburgh, University of Sheffield, Karlsruhe Institute of Technology, LIMSI, University of Amsterdam, Tilde). The systems are combined using RWTH’s system combination approach. The final submission shows an improvement of 1.0 BLEU compared to the best single system on newstest2016

Edinburgh Research Explorer

Publikationsserver der RWTH Aachen University

Biblio at Institute of Formal and Applied Linguistics

Wolverhampton Intellectual Repository and E-theses

Language Resources for Spanish - Spanish Sign Language (LSE) translation

Author: Garcia Adolfo
Lopez Ludeña Veronica
Martin Maganto Raquel
San Segundo Hernández Rubén
Sánchez David
Publication venue: E.T.S.I. Telecomunicación (UPM)
Publication date: 01/01/2010
Field of study

This paper describes the development of a Spanish Spanish Sign Language (LSE) translation system. Firstly, it describes the first Spanish Spanish Sign Language (LSE) parallel corpus focused on two specific domains: the renewal of the Identity Document and Driver’s License. This corpus includes more than 4,000 Spanish sentences (in these domains), their LSE translation and a video for each LSE sentence with the sign language representation. This corpus also contains more than 700 sign descriptions in several sign writing specifications. The translation system developed with this corpus consists of two modules: a Spanish into LSE translation module that is composed of a speech recognizer (for decoding the spoken utterance into a word sequence), a natural language translator (for converting a word sequence into a sequence of signs) and a 3D avatar animation module (for playing back the signs). The second module is a Spanish generator from LSE made up of a visual interface (for specifying a sequence of signs in sign writing), a language translator (for generating the sequence of words in Spanish) and a text to speech converter. For each language translation, the system uses three technologies: an example based strategy, a rule based translation method and a statistical translator

Archivo Digital UPM

A lightly supervised approach to detect stuttering in children's speech

Author: Alharbi S.
Brumfitt S.
Green P.
Hasan M.
Simons A.J.H.
Publication venue: 'International Speech Communication Association'
Publication date: 02/09/2018
Field of study

© 2018 International Speech Communication Association. All rights reserved. In speech pathology, new assistive technologies using ASR and machine learning approaches are being developed for detecting speech disorder events. Classically-trained ASR model tends to remove disfluencies from spoken utterances, due to its focus on producing clean and readable text output. However, diagnostic systems need to be able to track speech disfluencies, such as stuttering events, in order to determine the severity level of stuttering. To achieve this, ASR systems must be adapted to recognise full verbatim utterances, including pseudo-words and non-meaningful part-words. This work proposes a training regime to address this problem, and preserve a full verbatim output of stuttering speech. We use a lightly-supervised approach using task-oriented lattices to recognise the stuttering speech of children performing a standard reading task. This approach improved the WER by 27.8% relative to a baseline that uses word-lattices generated from the original prompt. The improved results preserved 63% of stuttering events (including sound, word, part-word and phrase repetition, and revision). This work also proposes a separate correction layer on top of the ASR that detects prolongation events (which are poorly recog-nised by the ASR). This increases the percentage of preserved stuttering events to 70%

Crossref

White Rose Research Online