Search CORE

19,367 research outputs found

Frame Based Single Channel Speech Separation using Sumary Autocorrelation Function

Author: Raghi E.R, Lekshmi M.S
Publication venue: 'Auricle Technologies, Pvt., Ltd.'
Publication date: 31/10/2015
Field of study

Single channel speech separation system is widely used in many applications. Pre-processing stage of Automatic speech recognition system, telecommunication system and the hearing aid design require the speech separation system to enhance the speech. This paper proposes a separation system that separates the dominant speech from the noisy environment, based on summary autocorrelation function (SACF) analysis pitch range estimation in the modulation frequency domain. Performance evaluation of the proposed system shows a better response compared to the existing methods

International Journal on Recent and Innovation Trends in Computing and Communication

Single channel speech separation in modulation frequency domain based on a novel pitch range estimation method

Author: A de Cheveigne
AS Bregman
D Talkin
DL Wang
G Hu
G Hu
G Hu
GJ Brown
J Barker
J Le Roux
J Tabrikian
JJ Sroka
L Atlas
M Buchler
M Wu
MH Radfar
MP Cooke
Q Li
R Drullman
RP Lippmann
S Dubnov
SM Schimmel
SM Schimmel
TW Lee
X Huang
Y Shao
Y Shao
Publication venue: 'Springer Science and Business Media LLC'
Publication date
Field of study

Crossref

Speech and crosstalk detection in multichannel audio

Author: Brown G.J.
Renals S.
Wan V.
Wrigley S.N.
Publication venue: 'Institute of Electrical and Electronics Engineers (IEEE)'
Publication date: 01/01/2005
Field of study

The analysis of scenarios in which a number of microphones record the activity of speakers, such as in a round-table meeting, presents a number of computational challenges. For example, if each participant wears a microphone, speech from both the microphone's wearer (local speech) and from other participants (crosstalk) is received. The recorded audio can be broadly classified in four ways: local speech, crosstalk plus local speech, crosstalk alone and silence. We describe two experiments related to the automatic classification of audio into these four classes. The first experiment attempted to optimize a set of acoustic features for use with a Gaussian mixture model (GMM) classifier. A large set of potential acoustic features were considered, some of which have been employed in previous studies. The best-performing features were found to be kurtosis, "fundamentalness," and cross-correlation metrics. The second experiment used these features to train an ergodic hidden Markov model classifier. Tests performed on a large corpus of recorded meetings show classification accuracies of up to 96%, and automatic speech recognition performance close to that obtained using ground truth segmentation

Crossref

Edinburgh Research Archive

Edinburgh Research Explorer

White Rose Research Online