Search CORE

15,735 research outputs found

A brief overview of speech enhancement with linear filtering

Author: Benesty Jacob
Chen Jingdong
Christensen Mads Græsbøll
Jensen Jesper Rindom
Publication venue: 'Springer Science and Business Media LLC'
Publication date: 01/01/2014
Field of study

Springer - Publisher Connector

VBN

A brief overview of speech enhancement with linear filtering

Author: B Kollmeier
GH Golub
H Huang
J Benesty
J Benesty
J Benesty
J Benesty
J Benesty
J Benesty
J Benesty
J Benesty
J Chen
J Chen
J Chen
J Freudenberger
JR Jensen
JR Jensen
K Hermus
LR Rabiner
M Dendrinos
M Souden
P Loizou
P Vary
R Martin
RC Hendriks
RJ McAulay
S Boll
S Doclo
S Srinivasan
SH Jensen
T Long
T Lotter
Y Ephraim
Y Ephraim
Y Hu
Publication venue: 'Springer Science and Business Media LLC'
Publication date
Field of study

Crossref

Deep Learning for Environmentally Robust Speech Recognition: An Overview of Recent Developments

Author: Geiger Jürgen
Jin Wenyu
Mousa Amr El-Desoky
Pohjalainen Jouni
Schuller Björn
Zhang Zixing
Publication venue
Publication date: 01/01/2018
Field of study

Eliminating the negative effect of non-stationary environmental noise is a long-standing research topic for automatic speech recognition that stills remains an important challenge. Data-driven supervised approaches, including ones based on deep neural networks, have recently emerged as potential alternatives to traditional unsupervised approaches and with sufficient training, can alleviate the shortcomings of the unsupervised methods in various real-life acoustic environments. In this light, we review recently developed, representative deep learning approaches for tackling non-stationary additive and convolutional degradation of speech with the aim of providing guidelines for those involved in the development of environmentally robust speech recognition systems. We separately discuss single- and multi-channel techniques developed for the front-end and back-end of speech recognition systems, as well as joint front-end and back-end training frameworks

arXiv.org e-Print Archive

OPUS Augsburg

Multi-Resolution Fully Convolutional Neural Networks for Monaural Audio Source Separation

Author: Grais Emad M.
Plumbley Mark D.
Ward Dominic
Wierstorf Hagen
Publication venue
Publication date: 28/10/2017
Field of study

In deep neural networks with convolutional layers, each layer typically has fixed-size/single-resolution receptive field (RF). Convolutional layers with a large RF capture global information from the input features, while layers with small RF size capture local details with high resolution from the input features. In this work, we introduce novel deep multi-resolution fully convolutional neural networks (MR-FCNN), where each layer has different RF sizes to extract multi-resolution features that capture the global and local details information from its input features. The proposed MR-FCNN is applied to separate a target audio source from a mixture of many audio sources. Experimental results show that using MR-FCNN improves the performance compared to feedforward deep neural networks (DNNs) and single resolution deep fully convolutional neural networks (FCNNs) on the audio source separation problem.Comment: arXiv admin note: text overlap with arXiv:1703.0801

arXiv.org e-Print Archive

University of Surrey

Surrey Research Insight

Online Monaural Speech Enhancement Using Delayed Subband LSTM

Author: Horaud Radu
Li Xiaofei
Publication venue
Publication date: 11/05/2020
Field of study

This paper proposes a delayed subband LSTM network for online monaural (single-channel) speech enhancement. The proposed method is developed in the short time Fourier transform (STFT) domain. Online processing requires frame-by-frame signal reception and processing. A paramount feature of the proposed method is that the same LSTM is used across frequencies, which drastically reduces the number of network parameters, the amount of training data and the computational burden. Training is performed in a subband manner: the input consists of one frequency, together with a few context frequencies. The network learns a speech-to-noise discriminative function relying on the signal stationarity and on the local spectral pattern, based on which it predicts a clean-speech mask at each frequency. To exploit future information, i.e. look-ahead, we propose an output-delayed subband architecture, which allows the unidirectional forward network to process a few future frames in addition to the current frame. We leverage the proposed method to participate to the DNS real-time speech enhancement challenge. Experiments with the DNS dataset show that the proposed method achieves better performance-measuring scores than the DNS baseline method, which learns the full-band spectra using a gated recurrent unit network.Comment: Paper submitted to Interspeech 202

arXiv.org e-Print Archive

Hal - Université Grenoble Alpes

INRIA a CCSD electronic archive server

Automatic Speech Recognition Using LP-DCTC/DCS Analysis Followed by Morphological Filtering

Author: Hix Penny
Publication venue: ODU Digital Commons
Publication date: 01/01/2006
Field of study

Front-end feature extraction techniques have long been a critical component in Automatic Speech Recognition (ASR). Nonlinear filtering techniques are becoming increasingly important in this application, and are often better than linear filters at removing noise without distorting speech features. However, design and analysis of nonlinear filters are more difficult than for linear filters. Mathematical morphology, which creates filters based on shape and size characteristics, is a design structure for nonlinear filters. These filters are limited to minimum and maximum operations that introduce a deterministic bias into filtered signals. This work develops filtering structures based on a mathematical morphology that utilizes the bias while emphasizing spectral peaks. The combination of peak emphasis via LP analysis with morphological filtering results in more noise robust speech recognition rates. To help understand the behavior of these pre-processing techniques the deterministic and statistical properties of the morphological filters are compared to the properties of feature extraction techniques that do not employ such algorithms. The robust behavior of these algorithms for automatic speech recognition in the presence of rapidly fluctuating speech signals with additive and convolutional noise is illustrated. Examples of these nonlinear feature extraction techniques are given using the Aurora 2.0 and Aurora 3.0 databases. Features are computed using LP analysis alone to emphasize peaks, morphological filtering alone, or a combination of the two approaches. Although absolute best results are normally obtained using a combination of the two methods, morphological filtering alone is nearly as effective and much more computationally efficient

Old Dominion University

Exploring the advantages of blind source separation in monitoring input respiratory impedance during apneic events

Author: De Keyser Robain
Ionescu Clara-Mihaela
Publication venue
Publication date: 01/01/2008
Field of study

Ghent University Academic Bibliography