Search CORE

983 research outputs found

The UPC Text-to-Speech System for Spanish and Catalan

Author: Bonafonte Cávez Antonio
Esquerra Llucià Ignasi
Febrer A
Rodríguez Fonollosa José Adrián
Vallverdú Bayés Sisco
Publication venue: 'The International Fiscal Association of Korea'
Publication date: 01/01/1998
Field of study

This paper summarizes the text-to-speech system that has been developed in the Speech Group of the Universitat Politècnica de Catalunya (UPC). The system is composed of a core and different interfaces so that it is compatible for research, for telephone applications (either CTI boards or standard ISDN PC cards supporting CAPI), and Windows applications developed using Microsoft SAPI. The paper reviews the system making emphasis in the parts of the system which are language dependent and which allow the reading of bilingual text (Spanish and Catalan). The paper also presents new approaches in prosodic modeling (segmental duration modeling) and generation of the database of speech segments, which have been introduced last year.Peer ReviewedPostprint (published version

UPCommons. Portal del coneixement obert de la UPC

An introduction to statistical parametric speech synthesis

Author: King Simon
Publication venue: 'Springer Science and Business Media LLC'
Publication date: 01/10/2011
Field of study

Edinburgh Research Explorer

Multilingual and Multimodal Corpus-Based Text-to-Speech System - PLATTOS -

Author: Izidor Mlakar
Matej Rojc
Publication venue: 'IntechOpen'
Publication date: 21/06/2011
Field of study

IntechOpen

Digital library of University of Maribor

Speech synthesis, Speech simulation and speech science

Author: Huckvale M
Publication venue
Publication date: 01/01/2002
Field of study

Speech synthesis research has been transformed in recent years through the exploitation of speech corpora - both for statistical modelling and as a source of signals for concatenative synthesis. This revolution in methodology and the new techniques it brings calls into question the received wisdom that better computer voice output will come from a better understanding of how humans produce speech. This paper discusses the relationship between this new technology of simulated speech and the traditional aims of speech science. The paper suggests that the goal of speech simulation frees engineers from inadequate linguistic and physiological descriptions of speech. But at the same time, it leaves speech scientists free to return to their proper goal of building a computational model of human speech production

UCL Discovery

Natural TTS Synthesis by Conditioning WaveNet on Mel Spectrogram Predictions

Author: Agiomyrgiannakis Yannis
Chen Zhifeng
Jaitly Navdeep
Pang Ruoming
Saurous Rif A.
Schuster Mike
Shen Jonathan
Skerry-Ryan RJ
Wang Yuxuan
Weiss Ron J.
Wu Yonghui
Yang Zongheng
Zhang Yu
Publication venue
Publication date: 15/02/2018
Field of study

This paper describes Tacotron 2, a neural network architecture for speech synthesis directly from text. The system is composed of a recurrent sequence-to-sequence feature prediction network that maps character embeddings to mel-scale spectrograms, followed by a modified WaveNet model acting as a vocoder to synthesize timedomain waveforms from those spectrograms. Our model achieves a mean opinion score (MOS) of

4.53

comparable to a MOS of

4.58

for professionally recorded speech. To validate our design choices, we present ablation studies of key components of our system and evaluate the impact of using mel spectrograms as the input to WaveNet instead of linguistic, duration, and

F_0

features. We further demonstrate that using a compact acoustic intermediate representation enables significant simplification of the WaveNet architecture.Comment: Accepted to ICASSP 201

arXiv.org e-Print Archive

Crossref

SMaTTS: standard malay text to speech system

Author: Ahmad Zakiah Hanim
Gunawan Teddy Surya
Khalifa Othman Omran
Publication venue: 'International Research Publication House'
Publication date: 01/01/2007
Field of study

This paper presents a rule-based text- to- speech (TTS) Synthesis System for Standard Malay, namely SMaTTS. The proposed system using sinusoidal method and some pre- recorded wave files in generating speech for the system. The use of phone database significantly decreases the amount of computer memory space used, thus making the system very light and embeddable. The overall system was comprised of two phases the Natural Language Processing (NLP) that consisted of the high-level processing of text analysis, phonetic analysis, text normalization and morphophonemic module. The module was designed specially for SM to overcome few problems in defining the rules for SM orthography system before it can be passed to the DSP module. The second phase is the Digital Signal Processing (DSP) which operated on the low-level process of the speech waveform generation. A developed an intelligible and adequately natural sounding formant-based speech synthesis system with a light and user-friendly Graphical User Interface (GUI) is introduced. A Standard Malay Language (SM) phoneme set and an inclusive set of phone database have been constructed carefully for this phone-based speech synthesizer. By applying the generative phonology, a comprehensive letter-to-sound (LTS) rules and a pronunciation lexicon have been invented for SMaTTS. As for the evaluation tests, a set of Diagnostic Rhyme Test (DRT) word list was compiled and several experiments have been performed to evaluate the quality of the synthesized speech by analyzing the Mean Opinion Score (MOS) obtained. The overall performance of the system as well as the room for improvements was thoroughly discussed

CiteSeerX

The International Islamic University Malaysia Repository

Study on phonetic context of Malay syllables towards the development of Malay speech synthesizer [TK7882.S65 H233 2007 f rb].

Author: Samsudin Nur Hana
Publication venue
Publication date: 01/01/2007
Field of study

Pensintesis sebutan Bahasa Melayu telah berkembang daripada teknik pensintesis berparameter (pemodelan penyebutan manusia dan pensintesis berdasarkan formant) kepada teknik pensintesis tidak berparameter (pensintesis sebutan berdasarkan pencantuman). Speech synthesizer has evolved from parametric speech synthesizer (articulatory and formant synthesizer) to non-parametric synthesizer (concatenative synthesizer). Recently, the concatenative speech synthesizer approach is moving towards corpusbased or unit selection technique

Repository@USM

Design and Implementation of Text To Speech Conversion for Visually Impaired People

Author: Isewon Itunuoluwa
Oladipupo O. O.
Oyelade O. J.
Publication venue
Publication date: 01/01/2012
Field of study

A Text-to-speech synthesizer is an application that converts text into spoken word, by analyzing and processing the text using Natural Language Processing (NLP) and then using Digital Signal Processing (DSP) technology to convert this processed text into synthesized speech representation of the text. Here, we developed a useful text-to-speech synthesizer in the form of a simple application that converts inputted text into synthesized speech and reads out to the user which can then be saved as an mp3.file. The development of a text to speech synthesizer will be of great help to people with visual impairment and make making through large volume of text easie

Covenant University Repository

Prosody in text-to-speech synthesis using fuzzy logic

Author: Williams Jonathan Brent
Publication venue: The Research Repository @ WVU
Publication date: 01/12/2005
Field of study

For over a thousand years, inventors, scientists and researchers have tried to reproduce human speech. Today, the quality of synthesized speech is not equivalent to the quality of real speech. Most research on speech synthesis focuses on improving the quality of the speech produced by Text-to-Speech (TTS) systems. The best TTS systems use unit selection-based concatenation to synthesize speech. However, this method is very timely and the speech database is very large. Diphone concatenated synthesized speech requires less memory, but sounds robotic. This thesis explores the use of fuzzy logic to make diphone concatenated speech sound more natural. A TTS is built using both neural networks and fuzzy logic. Text is converted into phonemes using neural networks. Fuzzy logic is used to control the fundamental frequency for three types of sentences. In conclusion, the fuzzy system produces f0 contours that make the diphone concatenated speech sound more natural

The Research Repository @ WVU (West Virginia University)