Search CORE

103 research outputs found

On Using Backpropagation for Speech Texture Generation and Voice Conversion

Author: Bengio Samy
Chorowski Jan
Saurous Rif A.
Weiss Ron J.
Publication venue
Publication date: 08/03/2018
Field of study

Inspired by recent work on neural network image generation which rely on backpropagation towards the network inputs, we present a proof-of-concept system for speech texture synthesis and voice conversion based on two mechanisms: approximate inversion of the representation learned by a speech recognition neural network, and on matching statistics of neuron activations between different source and target utterances. Similar to image texture synthesis and neural style transfer, the system works by optimizing a cost function with respect to the input waveform samples. To this end we use a differentiable mel-filterbank feature extraction pipeline and train a convolutional CTC speech recognition network. Our system is able to extract speaker characteristics from very limited amounts of target speaker data, as little as a few seconds, and can be used to generate realistic speech babble or reconstruct an utterance in a different voice.Comment: Accepted to ICASSP 201

arXiv.org e-Print Archive

Crossref

Evolution de la qualité des eaux du cours inférieur du fleuve de Nahr Ibrahim

Author: Aoun J.
Rif J.
Publication venue: Faculté des Sciences Agronomiques (Liban)
Publication date: 01/01/1998
Field of study

I-Revues

Natural TTS Synthesis by Conditioning WaveNet on Mel Spectrogram Predictions

Author: Agiomyrgiannakis Yannis
Chen Zhifeng
Jaitly Navdeep
Pang Ruoming
Saurous Rif A.
Schuster Mike
Shen Jonathan
Skerry-Ryan RJ
Wang Yuxuan
Weiss Ron J.
Wu Yonghui
Yang Zongheng
Zhang Yu
Publication venue
Publication date: 15/02/2018
Field of study

This paper describes Tacotron 2, a neural network architecture for speech synthesis directly from text. The system is composed of a recurrent sequence-to-sequence feature prediction network that maps character embeddings to mel-scale spectrograms, followed by a modified WaveNet model acting as a vocoder to synthesize timedomain waveforms from those spectrograms. Our model achieves a mean opinion score (MOS) of

4.53

comparable to a MOS of

4.58

for professionally recorded speech. To validate our design choices, we present ablation studies of key components of our system and evaluate the impact of using mel spectrograms as the input to WaveNet instead of linguistic, duration, and

F_0

features. We further demonstrate that using a compact acoustic intermediate representation enables significant simplification of the WaveNet architecture.Comment: Accepted to ICASSP 201

arXiv.org e-Print Archive

Crossref

Tacotron: Towards End-to-End Speech Synthesis

Author: Agiomyrgiannakis Yannis
Bengio Samy
Chen Zhifeng
Clark Rob
Jaitly Navdeep
Le Quoc
Saurous Rif A.
Skerry-Ryan RJ
Stanton Daisy
Wang Yuxuan
Weiss Ron J.
Wu Yonghui
Xiao Ying
Yang Zongheng
Publication venue
Publication date: 06/04/2017
Field of study

A text-to-speech synthesis system typically consists of multiple stages, such as a text analysis frontend, an acoustic model and an audio synthesis module. Building these components often requires extensive domain expertise and may contain brittle design choices. In this paper, we present Tacotron, an end-to-end generative text-to-speech model that synthesizes speech directly from characters. Given pairs, the model can be trained completely from scratch with random initialization. We present several key techniques to make the sequence-to-sequence framework perform well for this challenging task. Tacotron achieves a 3.82 subjective 5-scale mean opinion score on US English, outperforming a production parametric system in terms of naturalness. In addition, since Tacotron generates speech at the frame level, it's substantially faster than sample-level autoregressive methods.Comment: Submitted to Interspeech 2017. v2 changed paper title to be consistent with our conference submission (no content change other than typo fixes

arXiv.org e-Print Archive

Crossref

CNN Architectures for Large-Scale Audio Classification

Author: Chaudhuri Sourish
Ellis Daniel P. W.
Gemmeke Jort F.
Hershey Shawn
Jansen Aren
Moore R. Channing
Plakal Manoj
Platt Devin
Saurous Rif A.
Seybold Bryan
Slaney Malcolm
Weiss Ron J.
Wilson Kevin
Publication venue
Publication date: 10/01/2017
Field of study

Convolutional Neural Networks (CNNs) have proven very effective in image classification and show promise for audio. We use various CNN architectures to classify the soundtracks of a dataset of 70M training videos (5.24 million hours) with 30,871 video-level labels. We examine fully connected Deep Neural Networks (DNNs), AlexNet [1], VGG [2], Inception [3], and ResNet [4]. We investigate varying the size of both training set and label vocabulary, finding that analogs of the CNNs used in image classification do well on our audio classification task, and larger training and label sets help up to a point. A model using embeddings from these classifiers does much better than raw features on the Audio Set [5] Acoustic Event Detection (AED) classification task.Comment: Accepted for publication at ICASSP 2017 Changes: Added definitions of mAP, AUC, and d-prime. Updated mAP/AUC/d-prime numbers for Audio Set based on changes of latest Audio Set revision. Changed wording to fit 4 page limit with new addition

arXiv.org e-Print Archive

Crossref

Rock magnetic signature of the Middle Eocene Climatic Optimum (MECO) event in different oceanic basins

Author: Catanzariti R
Chang L
Coccioni R
Florindo F
Frontalini F
Giorgioni M
Iacoviello F
Jovane L
Roberts AP
Rodelli D
Savian J
Sprovieri M
Trindade RIF
Publication venue: LATINMAG LETTERS
Publication date: 15/10/2015
Field of study

The Middle Eocene Climatic Optimum (MECO) event at ~40 Ma was a greenhouse warming which indicates an abrupt reversal in long-term cooling through the middle Eocene. Here, we present environmental and rock magnetic data from sedimentary successions from the Indian Ocean (ODP Hole 711A) and eastern NeoTethys (Monte Cagnero section - MCA). The high-resolution environmental magnetism record obtained for MCA section shows an interval of increase of magnetic parameters comprising the MECO peak. A relative increase in eutrophic nannofossil taxa spans the culmination of the MECO warming and its aftermath and coincides with a positive carbon isotope excursion, and a peak in magnetite and hematite/goethite concentrations. The magnetite peak reflects the appearance of magnetofossils, while the hematite/goethite apex are attributed to an enhanced detrital mineral contribution, likely related to aeolian dust transported from the continent adjacent to the Neo-Tethys Ocean during a drier, more seasonal MECO climate. Seasurface iron fertilization is inferred to have stimulated high phytoplankton productivity, increasing organic carbon export to the seafloor and promoting enhanced biomineralization of magnetotactic bacteria, which are preserved as magnetofossils during the warmest periods of the MECO event. Environmental magnetic parameters show the same behavior for ODP Hole 711A. We speculate that iron fertilization promoted by aeolian hematite during the MECO event has contributed significantly to increase the primary productivity in the oceans. The widespread occurrence of magnetofossils in other warming periods suggests a common mechanism linking climate warming and enhancement of magnetosome production and preservation

UCL Discovery

Helping Parents Make Sense of Video Game Addiction

Author: A Brus
A Gade
A Parkes
A Rooij Van
AK Przybylski
AK Przybylski
AM Bean
American Psychiatric Association
Blizzard Entertainment
D Kardefelt-Winther
D Kardefelt-Winther
D Kardefelt-Winther
E Aarseth
I Granic
J Billieux
J Charlton
J Enevold
J Phyfer
K Durkin
KC Berridge
M Scharkow
MJ George
MJ Koepp
N Weinstein
NM Petry
R Cover
R Kowert
RIF Brown
RKL Nielsen
S Livingstone
S Turkle
UNICEF
W Glasser
Publication venue: 'Springer Science and Business Media LLC'
Publication date: 22/08/2018
Field of study

Crossref

The IT University of Copenhagen's Repository

Neural Correlates of Auditory Perceptual Awareness under Informational Masking

Author: Alain
Alexander Gutschalk
Andrew J Oxenham
Androulidakis
Bidet-Caulet
Bledowski
Bregman
Brungart
Christophe Micheyl
Corbetta
Delgutte
Desimone
Durlach
Efron
Evans
Fishman
Galaburda
Galambos
Galambos
Godey
Gutschalk
Gutschalk
Gutschalk
Hari
Hillyard
Hillyard
J&auml
Kanwisher
Kidd
Kidd
Lavie
Leonard
Liegeois-Chauvel
Lins
Micheyl
Micheyl
Moore
N&auml
Neff
Neff
Neff
Oates
Okamoto
Overath
Pantev
Parasuraman
Picton
Pockett
Rademacher
Richards
Rif
Rivier
Romani
Scherg
Shera
Snyder
Steinschneider
Timothy D Griffiths
Watson
Wegel
Woldorff
Publication venue: Public Library of Science
Publication date: 01/06/2008
Field of study

Our ability to detect target sounds in complex acoustic backgrounds is often limited not by the ear's resolution, but by the brain's information-processing capacity. The neural mechanisms and loci of this “informational masking” are unknown. We combined magnetoencephalography with simultaneous behavioral measures in humans to investigate neural correlates of informational masking and auditory perceptual awareness in the auditory cortex. Cortical responses were sorted according to whether or not target sounds were detected by the listener in a complex, randomly varying multi-tone background known to produce informational masking. Detected target sounds elicited a prominent, long-latency response (50–250 ms), whereas undetected targets did not. In contrast, both detected and undetected targets produced equally robust auditory middle-latency, steady-state responses, presumably from the primary auditory cortex. These findings indicate that neural correlates of auditory awareness in informational masking emerge between early and late stages of processing within the auditory cortex

Public Library of Science (PLOS)

Crossref

Directory of Open Access Journals

PubMed Central