Search CORE

9 research outputs found

Visual Speech Enhancement

Author: Gabbay Aviv
Peleg Shmuel
Shamir Asaph
Publication venue
Publication date: 13/06/2018
Field of study

When video is shot in noisy environment, the voice of a speaker seen in the video can be enhanced using the visible mouth movements, reducing background noise. While most existing methods use audio-only inputs, improved performance is obtained with our visual speech enhancement, based on an audio-visual neural network. We include in the training data videos to which we added the voice of the target speaker as background noise. Since the audio input is not sufficient to separate the voice of a speaker from his own voice, the trained model better exploits the visual input and generalizes well to different noise types. The proposed model outperforms prior audio visual methods on two public lipreading datasets. It is also the first to be demonstrated on a dataset not designed for lipreading, such as the weekly addresses of Barack Obama.Comment: Accepted to Interspeech 2018. Supplementary video: https://www.youtube.com/watch?v=nyYarDGpcY

arXiv.org e-Print Archive

Crossref

SEGAN: Speech Enhancement Generative Adversarial Network

Author: Bonafonte Antonio
Pascual Santiago
Serrà Joan
Publication venue
Publication date: 09/06/2017
Field of study

Current speech enhancement techniques operate on the spectral domain and/or exploit some higher-level feature. The majority of them tackle a limited number of noise conditions and rely on first-order statistics. To circumvent these issues, deep networks are being increasingly used, thanks to their ability to learn complex functions from large example sets. In this work, we propose the use of generative adversarial networks for speech enhancement. In contrast to current techniques, we operate at the waveform level, training the model end-to-end, and incorporate 28 speakers and 40 different noise conditions into the same model, such that model parameters are shared across them. We evaluate the proposed model using an independent, unseen test set with two speakers and 20 alternative noise conditions. The enhanced samples confirm the viability of the proposed model, and both objective and subjective evaluations confirm the effectiveness of it. With that, we open the exploration of generative architectures for speech enhancement, which may progressively incorporate further speech-centric design choices to improve their performance.Comment: 5 pages, 4 figures, accepted in INTERSPEECH 201

arXiv.org e-Print Archive

Crossref

Studies on noise robust automatic speech recognition

Author: Kurimo Mikko
Palomäki Kalle J.
Remes Ulpu
Publication venue: Teknillinen korkeakoulu
Publication date: 01/01/2009
Field of study

Noise in everyday acoustic environments such as cars, traffic environments, and cafeterias remains one of the main challenges in automatic speech recognition (ASR). As a research theme, it has received wide attention in conferences and scientific journals focused on speech technology. This article collection reviews both the classic and novel approaches suggested for noise robust ASR. The articles are literature reviews written for the spring 2009 seminar course on noise robust automatic speech recognition (course code T-61.6060) held at TKK

Aaltodoc Publication Archive

Improved sequential and batch learning in neural networks using the tangent plane algorithm

Author: May Paul
Publication venue
Publication date: 01/01/2012
Field of study

The principal aim of this research is to investigate and develop improved sequential and batch learning algorithms based upon the tangent plane algorithm for artificial neural networks. A secondary aim is to apply the newly developed algorithms to multi-category cancer classification problems in the bio-informatics area, which involves the study of dna or protein sequences, macro-molecular structures, and gene expressions

University of Bolton Institutional Repository (UBIR)

University of Bolton Institutional Repository

Efficient, end-to-end and self-supervised methods for speech processing and generation

Author: Pascual de la Puente Santiago
Publication venue: Universitat Politècnica de Catalunya
Publication date: 31/01/2020
Field of study

Deep learning has affected the speech processing and generation fields in many directions. First, end-to-end architectures allow the direct injection and synthesis of waveform samples. Secondly, the exploration of efficient solutions allow to implement these systems in computationally restricted environments, like smartphones. Finally, the latest trends exploit audio-visual data with least supervision. In this thesis these three directions are explored. Firstly, we propose the use of recent pseudo-recurrent structures, like self-attention models and quasi-recurrent networks, to build acoustic models for text-to-speech. The proposed system, QLAD, turns out to synthesize faster on CPU and GPU than its recurrent counterpart whilst preserving the good synthesis quality level, which is competitive with state of the art vocoder-based models. Then, a generative adversarial network is proposed for speech enhancement, named SEGAN. This model works as a speech-to-speech conversion system in time-domain, where a single inference operation is needed for all samples to operate through a fully convolutional structure. This implies an increment in modeling efficiency with respect to other existing models, which are auto-regressive and also work in time-domain. SEGAN achieves prominent results in noise supression and preservation of speech naturalness and intelligibility when compared to the other classic and deep regression based systems. We also show that SEGAN is efficient in transferring its operations to new languages and noises. A SEGAN trained for English performs similarly to this language on Catalan and Korean with only 24 seconds of adaptation data. Finally, we unveil the generative capacity of the model to recover signals from several distortions. We hence propose the concept of generalized speech enhancement. First, the model proofs to be effective to recover voiced speech from whispered one. Then the model is scaled up to solve other distortions that require a recomposition of damaged parts of the signal, like extending the bandwidth or recovering lost temporal sections, among others. The model improves by including additional acoustic losses in a multi-task setup to impose a relevant perceptual weighting on the generated result. Moreover, a two-step training schedule is also proposed to stabilize the adversarial training after the addition of such losses, and both components boost SEGAN's performance across distortions.Finally, we propose a problem-agnostic speech encoder, named PASE, together with the framework to train it. PASE is a fully convolutional network that yields compact representations from speech waveforms. These representations contain abstract information like the speaker identity, the prosodic features or the spoken contents. A self-supervised framework is also proposed to train this encoder, which suposes a new step towards unsupervised learning for speech processing. Once the encoder is trained, it can be exported to solve different tasks that require speech as input. We first explore the performance of PASE codes to solve speaker recognition, emotion recognition and speech recognition. PASE works competitively well compared to well-designed classic features in these tasks, specially after some supervised adaptation. Finally, PASE also provides good descriptors of identity for multi-speaker modeling in text-to-speech, which is advantageous to model novel identities without retraining the model.L'aprenentatge profund ha afectat els camps de processament i generació de la parla en vàries direccions. Primer, les arquitectures fi-a-fi permeten la injecció i síntesi de mostres temporals directament. D'altra banda, amb l'exploració de solucions eficients permet l'aplicació d'aquests sistemes en entorns de computació restringida, com els telèfons intel·ligents. Finalment, les darreres tendències exploren les dades d'àudio i veu per derivar-ne representacions amb la mínima supervisió. En aquesta tesi precisament s'exploren aquestes tres direccions. Primer de tot, es proposa l'ús d'estructures pseudo-recurrents recents, com els models d’auto atenció i les xarxes quasi-recurrents, per a construir models acústics text-a-veu. Així, el sistema QLAD proposat en aquest treball sintetitza més ràpid en CPU i GPU que el seu homòleg recurrent, preservant el mateix nivell de qualitat de síntesi, competitiu amb l'estat de l'art en models basats en vocoder. A continuació es proposa un model de xarxa adversària generativa per a millora de veu, anomenat SEGAN. Aquest model fa conversions de veu-a-veu en temps amb una sola operació d'inferència sobre una estructura purament convolucional. Això implica un increment en l'eficiència respecte altres models existents auto regressius i que també treballen en el domini temporal. La SEGAN aconsegueix resultats prominents d'extracció de soroll i preservació de la naturalitat i la intel·ligibilitat de la veu comparat amb altres sistemes clàssics i models regressius basats en xarxes neuronals profundes en espectre. També es demostra que la SEGAN és eficient transferint les seves operacions a nous llenguatges i sorolls. Així, un model SEGAN entrenat en Anglès aconsegueix un rendiment comparable a aquesta llengua quan el transferim al català o al coreà amb només 24 segons de dades d'adaptació. Finalment, explorem l'ús de tota la capacitat generativa del model i l’apliquem a recuperació de senyals de veu malmeses per vàries distorsions severes. Això ho anomenem millora de la parla generalitzada. Primer, el model demostra ser efectiu per a la tasca de recuperació de senyal sonoritzat a partir de senyal xiuxiuejat. Posteriorment, el model escala a poder resoldre altres distorsions que requereixen una reconstrucció de parts del senyal que s’han malmès, com extensió d’ample de banda i recuperació de seccions temporals perdudes, entre d’altres. En aquesta última aplicació del model, el fet d’incloure funcions de pèrdua acústicament rellevants incrementa la naturalitat del resultat final, en una estructura multi-tasca que prediu característiques acústiques a la sortida de la xarxa discriminadora de la nostra GAN. També es proposa fer un entrenament en dues etapes del sistema SEGAN, el qual mostra un increment significatiu de l’equilibri en la sinèrgia adversària i la qualitat generada finalment després d’afegir les funcions acústiques. Finalment, proposem un codificador de veu agnòstic al problema, anomenat PASE, juntament amb el conjunt d’eines per entrenar-lo. El PASE és un sistema purament convolucional que crea representacions compactes de trames de veu. Aquestes representacions contenen informació abstracta com identitat del parlant, les característiques prosòdiques i els continguts lingüístics. També es proposa un entorn auto-supervisat multi-tasca per tal d’entrenar aquest sistema, el qual suposa un avenç en el terreny de l’aprenentatge no supervisat en l’àmbit del processament de la parla. Una vegada el codificador esta entrenat, es pot exportar per a solventar diferents tasques que requereixin tenir senyals de veu a l’entrada. Primer explorem el rendiment d’aquest codificador per a solventar tasques de reconeixement del parlant, de l’emoció i de la parla, mostrant-se efectiu especialment si s’ajusta la representació de manera supervisada amb un conjunt de dades d’adaptació.Postprint (published version

UPCommons. Portal del coneixement obert de la UPC

Efficient, end-to-end and self-supervised methods for speech processing and generation

Author: Pascual De La Puente Santiago
Publication venue: Universitat Politècnica de Catalunya
Publication date: 01/01/2020
Field of study

LAReferencia - Red Federada de Repositorios Institucionales de Publicaciones Científicas Latinoamericanas

UPCommons. Portal del coneixement obert de la UPC

Tesis Doctorals en Xarxa

Estructuras de metadatos para un mejor uso de la información en alimentación animal

Author: Maroto Molina Francisco
Publication venue: Universidad de Córdoba, Servicio de Publicaciones
Publication date: 01/01/2013
Field of study

Para un uso eficiente de los alimentos es necesario conocer en profundidad tanto las necesidades de los animales como las características de los alimentos. Respecto a estas últimas, los datos sobre la composición química y el valor nutritivo de los alimentos se han obtenido de forma sistemática en los laboratorios de nutrición animal durante los últimos 200 años (Gizzi y Givens, 2004). Sin embargo, la mayoría de estos datos se utilizan con un propósito único, ya sea el control de calidad o la producción de resultados científicos, obviando su valor residual cuando se analizan conjuntamente. Desde principios del siglo XX, parte de esta información se ha recogido en tablas, pero éstas presentan algunas limitaciones, como el tamaño reducido y la estaticidad. Para superar dichas limitaciones surgen las bases de datos de alimentos. El Servicio de Información sobre Alimentos (SIA) de la Universidad de Córdoba lleva años trabajando en la construcción de este tipo de bases de datos (Gómez Cabrera et al., 2003), pero se ha encontrado con algunas dificultades relacionadas con la gestión y análisis de la información acumulada. La búsqueda de soluciones a dichos problemas es el punto de partida de la presente Tesis Doctoral. 2. Contenido de la investigación Los datos acumulados en el SIA carecían de la información accesoria necesaria para una adecuada interpretación y uso. Para solucionarlo se ha diseñado una estructura de metadatación adaptada a las necesidades del registro diario de información en los laboratorios. Por otro lado, se han diseñado sistemáticas de denominación y lenguajes controlados para los metadatos de cara a evitar la heterogeneidad de los descriptores. Respecto a la fase de análisis de la información, se había detectado la importancia del pre-procesamiento. En la presente Tesis Doctoral se ha estudiado el comportamiento, respecto a las externalidades o outputs más habituales de las bases de datos de alimentos, de diferentes técnicas para la integración de datos diversos, la búsqueda de repeticiones, la detección de anómalos o outliers y la gestión de los vacíos de información o missing data. Se han estudiado algoritmos uni- y multivariantes, así como aproximaciones globales y locales a los aspectos citados. 3. Conclusión Se concluye que las bases de datos de alimentos construidas en base a estructuras de metadatos son una gran opción para compartir resultados de investigación (data sharing) y para controlar la heterogeneidad típica de los datos sobre alimentos para animales. El pre-procesamiento de la información, en especial la detección de outliers y el manejo de missing data, se muestra como un paso esencial, siendo los algoritmos más adecuados en cada caso función de las características de la base de datos y del tipo de análisis que se quiere llevar a cabo. Además, pese a que ambos aspectos suelen ser vistos como un problema, su estudio permite obtener información cuali- y cuantitativa muy valiosa

Repositorio Institucional de la Universidad de Córdoba