Speaker adaptation and the evaluation of speaker                   similarity in the EMIME speech-to-speech translation                   project

Byrne, William; Dines, John; Garner, Philip N.; Gibson, Matthew; Guan, Yong; Hirsimäki, Teemu; Karhila, Reima; King, Simon; Kurimo, Mikko; Liang, Hui; Oura, Keiichiro; Saheer, Lakshmi; Shannon, Matt; Shiota, Sayaka; Tian, Jilei; Tokuda, Keiichi; Wester, Mirjam; Wu, Yi-Jian; Yamagishi, Junichi

Speaker adaptation and the evaluation of speaker similarity in the EMIME speech-to-speech translation project

Authors: William Byrne
John Dines
Philip N. Garner
Matthew Gibson
Yong Guan
Teemu Hirsimäki
Reima Karhila
Simon King
Mikko Kurimo
Hui Liang
Keiichiro Oura
Lakshmi Saheer
Matt Shannon
Sayaka Shiota
Jilei Tian
Keiichi Tokuda
Mirjam Wester
Yi-Jian Wu
Junichi Yamagishi
Publication date: 1 January 2010
Publisher: 7th ISCA Speech Synthesis Workshop

Abstract

This paper provides an overview of speaker adaptation research carried out in the EMIME speech-to-speech translation (S2ST) project. We focus on how speaker adaptation transforms can be learned from speech in one language and applied to the acoustic models of another language. The adaptation is transferred across languages and/or from recognition models to synthesis models. The various approaches investigated can all be viewed as a process in which a mapping is defined in terms of either acoustic model states or linguistic units. The mapping is used to transfer either speech data or adaptation transforms between the two models. Because the success of speaker adaptation in text-to-speech synthesis is measured by judging speaker similarity, we also discuss issues concerning evaluation of speaker similarity in an S2ST scenario