Search CORE

58 research outputs found

The JHU Parallel Corpus Filtering Systems for WMT 2018

Author: Khayrallah Huda
Koehn Philipp
Xu Hainan
Publication venue: 'Association for Computational Linguistics (ACL)'
Publication date: 01/01/2018
Field of study

Crossref

Edinburgh Research Explorer

Findings of the 2019 Conference on Machine Translation (WMT19)

Author: Barrault Loïc
Bojar Ondřej
Costa-Jussà Marta R.
Federmann Christian
Fishel Mark
Graham Yvette
Publication venue: 'Association for Computational Linguistics (ACL)'
Publication date: 01/08/2019
Field of study

This paper presents the results of the premier shared task organized alongside the Conference on Machine Translation (WMT) 2019. Participants were asked to build machine translation systems for any of 18 language pairs, to be evaluated on a test set of news stories. The main metric for this task is human judgment of translation quality. The task was also opened up to additional test suites to probe specific aspects of translation

Irish Universities

DCU Online Research Access Service

Evaluating conjunction disambiguation on English-to-German and French-to-German WMT 2019 translation hypotheses

Author: Popović Maja
Publication venue: 'Association for Computational Linguistics (ACL)'
Publication date: 01/01/2019
Field of study

We present a test set for evaluating an MT system’s capability to translate ambiguous conjunctions depending on the sentence structure. We concentrate on the English conjunction ”but” and its French equivalent ”mais” which can be translated into two different German conjunctions. We evaluate all English-to-German and French-to-German submissions to the WMT 2019 shared translation task. The evaluation is done mainly automatically, with additional fast manual inspection of unclear cases. All systems almost perfectly recognise the ta-get conjunction ”aber”, whereas accuracies fo rthe other target conjunction ”sondern” range from 78% to 97%, and the errors are mostly caused by replacing it with the alternative cojjunction ”aber”. The best performing system for both language pairs is a multilingual Transformer TartuNLP system trained on all WMT2019 language pairs which use the Latin script, indicating that the multilingual approach is beneficial for conjunction disambiguation. As for other system features, such as using synthetic back-translated data, context-aware, hybrid, etc., no particular (dis)advantages can be observed. Qualitative manual inspection of translation hypotheses shown that highly ranked systems generally produce translations with high adequacy and fluency, meaning that these systems are not only capable of capturing the right conjunction whereas the rest of the translation hypothesis is poor. On the other hand, the low ranked systems generally exhibit lower fluency and poor adequacy

Crossref

Irish Universities

DCU Online Research Access Service

The University of Edinburgh's English-German and English-Hausa Submissions to the WMT21 News Translation Task

Author: Birch Alexandra
Bogoychev Nikolay
Burchell Laurie
Chen Pinzhen
Germann Ulrich
Heafield Kenneth
Helcl Jindřich
Miceli Barone Antonio Valerio
Waldendorf Jonas
Publication venue
Publication date: 01/11/2021
Field of study

This paper presents the University of Edinburgh's constrained submissions of English-German and English-Hausa systems to the WMT 2021 shared task on news translation. We build En-De systems in three stages: corpus filtering, back-translation, and fine-tuning. For En-Ha we use an iterative back-translation approach on top of pre-trained En-De models and investigate vocabulary embedding mapping

Edinburgh Research Explorer

On the Evaluation of Semantic Phenomena in Neural Machine Translation Using Natural Language Inference

Author: Belinkov Yonatan
Glass James
Poliak Adam
Van Durme Benjamin
Publication venue
Publication date: 01/01/2018
Field of study

We propose a process for investigating the extent to which sentence representations arising from neural machine translation (NMT) systems encode distinct semantic phenomena. We use these representations as features to train a natural language inference (NLI) classifier based on datasets recast from existing semantic annotations. In applying this process to a representative NMT system, we find its encoder appears most suited to supporting inferences at the syntax-semantics interface, as compared to anaphora resolution requiring world-knowledge. We conclude with a discussion on the merits and potential deficiencies of the existing process, and how it may be improved and extended as a broader framework for evaluating semantic coverage.Comment: To be presented at NAACL 2018 - 11 page

arXiv.org e-Print Archive

Crossref