Search CORE

408 research outputs found

The Research of Noise-Robust Speech Recognition Based on Frequency Warping Wavelet

Author: Wenjun Meng
Xueying Zhang
Publication venue: 'IntechOpen'
Publication date: 01/06/2007
Field of study

Wavelet transforms for non-uniform speech recognition

Author: Javier L
Lleida E
Martí J
Nadeu Camprubí Climent
Publication venue: 'Institute of Electrical and Electronics Engineers (IEEE)'
Publication date: 01/01/1996
Field of study

An algorithm for nonuniform speech segmentation and its application in speech recognition systems is presented. A method based on the Modulated Gaussian Wavelet Transform based Speech Analyser (MGWTSA) and the subsequent parametrization block is used to transform a uniform signal into a set of nonuniformly separated frames, with the accurate information being fed into a speech recognition system. The algorithm needs a frame characterizing the signal where necessary, trying to reduce the number of frames per signal as much as possible, without an appreciable reduction in the recognition rate of the system.Peer ReviewedPostprint (published version

UPCommons. Portal del coneixement obert de la UPC

Some Commonly Used Speech Feature Extraction Algorithms

Author: Alim Sabur Ajibola
Rashid Nahrul Khair Alang
Publication venue: 'IntechOpen'
Publication date: 12/12/2018
Field of study

Speech is a complex naturally acquired human motor ability. It is characterized in adults with the production of about 14 different sounds per second via the harmonized actions of roughly 100 muscles. Speaker recognition is the capability of a software or hardware to receive speech signal, identify the speaker present in the speech signal and recognize the speaker afterwards. Feature extraction is accomplished by changing the speech waveform to a form of parametric representation at a relatively minimized data rate for subsequent processing and analysis. Therefore, acceptable classification is derived from excellent and quality features. Mel Frequency Cepstral Coefficients (MFCC), Linear Prediction Coefficients (LPC), Linear Prediction Cepstral Coefficients (LPCC), Line Spectral Frequencies (LSF), Discrete Wavelet Transform (DWT) and Perceptual Linear Prediction (PLP) are the speech feature extraction techniques that were discussed in these chapter. These methods have been tested in a wide variety of applications, giving them high level of reliability and acceptability. Researchers have made several modifications to the above discussed techniques to make them less susceptible to noise, more robust and consume less time. In conclusion, none of the methods is superior to the other, the area of application would determine which method to select

IntechOpen

Crossref

A Framework for Bioacoustic Vocalization Analysis Using Hidden Markov Models

Author: Ren Yao
Johnson Michael T
Clemins Patrick J.
Darre Michael
Glaeser Sharon Stuart
Osiejuk Tomasz S.
Out-Nyarko Ebenezer
Publication venue: e-Publications@Marquette
Publication date: 19/07/1999
Field of study

Using Hidden Markov Models (HMMs) as a recognition framework for automatic classification of animal vocalizations has a number of benefits, including the ability to handle duration variability through nonlinear time alignment, the ability to incorporate complex language or recognition constraints, and easy extendibility to continuous recognition and detection domains. In this work, we apply HMMs to several different species and bioacoustic tasks using generalized spectral features that can be easily adjusted across species and HMM network topologies suited to each task. This experimental work includes a simple call type classification task using one HMM per vocalization for repertoire analysis of Asian elephants, a language-constrained song recognition task using syllable models as base units for ortolan bunting vocalizations, and a stress stimulus differentiation task in poultry vocalizations using a non-sequential model via a one-state HMM with Gaussian mixtures. Results show strong performance across all tasks and illustrate the flexibility of the HMM framework for a variety of species, vocalization types, and analysis tasks

epublications@Marquette

University of Sheffield Library Digital Collections

A Framework for Bioacoustic Vocalization Analysis Using Hidden Markov Models

Author: Clemins Patrick J.
Darre Michael
Glaeser Sharon Stuart
Johnson Michael T
Osiejuk Tomasz S.
Out-Nyarko Ebenezer
Ren Yao
Publication venue: e-Publications@Marquette
Publication date: 01/11/2009
Field of study

Multidisciplinary Digital Publishing Institute

epublications@Marquette

Directory of Open Access Journals

The Teager-Kaiser Energy Cepstral Coefficients as an Effective Structural Health Monitoring Tool

Author: Betti Raimondo
Ceravolo Rosario
Civera Marco
Ferraris Matteo
Surace Cecilia
Publication venue: 'MDPI AG'
Publication date: 23/11/2019
Field of study

Recently, features and techniques from speech processing have started to gain increasing attention in the Structural Health Monitoring (SHM) community, in the context of vibration analysis. In particular, the Cepstral Coefficients (CCs) proved to be apt in discerning the response of a damaged structure with respect to a given undamaged baseline. Previous works relied on the Mel-Frequency Cepstral Coefficients (MFCCs). This approach, while efficient and still very common in applications, such as speech and speaker recognition, has been followed by other more advanced and competitive techniques for the same aims. The Teager-Kaiser Energy Cepstral Coefficients (TECCs) is one of these alternatives. These features are very closely related to MFCCs, but provide interesting and useful additional values, such as e.g., improved robustness with respect to noise. The goal of this paper is to introduce the use of TECCs for damage detection purposes, by highlighting their competitiveness with closely related features. Promising results from both numerical and experimental data were obtained

Multidisciplinary Digital Publishing Institute

PORTO@iris (Publications Open Repository TOrino - Politecnico di Torino)