Search CORE

223 research outputs found

A time-frequency approach to blind separation of under-determined mixture of sources

Author: Kawamoto M.
Mansour Ali
Puntonet C.
Publication venue: IASTED
Publication date: 01/01/2003
Field of study

espace@Curtin

Proceedings of the Second Conference of Students of Systematic Musicology

Author: Naveda Luiz Alberto
Publication venue: 'Ghent University'
Publication date: 01/01/2009
Field of study

Ghent University Academic Bibliography

Studies on noise robust automatic speech recognition

Author: Kurimo Mikko
Palomäki Kalle J.
Remes Ulpu
Publication venue: Teknillinen korkeakoulu
Publication date: 01/01/2009
Field of study

Noise in everyday acoustic environments such as cars, traffic environments, and cafeterias remains one of the main challenges in automatic speech recognition (ASR). As a research theme, it has received wide attention in conferences and scientific journals focused on speech technology. This article collection reviews both the classic and novel approaches suggested for noise robust ASR. The articles are literature reviews written for the spring 2009 seminar course on noise robust automatic speech recognition (course code T-61.6060) held at TKK

Aaltodoc Publication Archive

Quantitative assessment of spatial sound distortion by the semi-ideal recording point of a hear-through device

Author: Christensen Flemming
Hammershøi Dorte
Hoffmann Pablo F.
Publication venue: 'Acoustical Society of America (ASA)'
Publication date: 01/01/2013
Field of study

Crossref

VBN

Spatial dissection of a soundfield using spherical harmonic decomposition

Author: Fahim Abdullah
Publication venue
Publication date: 01/01/2020
Field of study

A real-world soundfield is often contributed by multiple desired and undesired sound sources. The performance of many acoustic systems such as automatic speech recognition, audio surveillance, and teleconference relies on its ability to extract the desired sound components in such a mixed environment. The existing solutions to the above problem are constrained by various fundamental limitations and require to enforce different priors depending on the acoustic condition such as reverberation and spatial distribution of sound sources. With the growing emphasis and integration of audio applications in diverse technologies such as smart home and virtual reality appliances, it is imperative to advance the source separation technology in order to overcome the limitations of the traditional approaches. To that end, we exploit the harmonic decomposition model to dissect a mixed soundfield into its underlying desired and undesired components based on source and signal characteristics. By analysing the spatial projection of a soundfield, we achieve multiple outcomes such as (i) soundfield separation with respect to distinct source regions, (ii) source separation in a mixed soundfield using modal coherence model, and (iii) direction of arrival (DOA) estimation of multiple overlapping sound sources through pattern recognition of the modal coherence of a soundfield. We first employ an array of higher order microphones for soundfield separation in order to reduce hardware requirement and implementation complexity. Subsequently, we develop novel mathematical models for modal coherence of noisy and reverberant soundfields that facilitate convenient ways for estimating DOA and power spectral densities leading to robust source separation algorithms. The modal domain approach to the soundfield/source separation allows us to circumvent several practical limitations of the existing techniques and enhance the performance and robustness of the system. The proposed methods are presented with several practical applications and performance evaluations using simulated and real-life dataset

The Australian National University

The impact of voice phonetics on audience's perception in the context of activist video

Author: Nina Holesova
Publication venue
Publication date: 23/07/2014
Field of study

Repositório Aberto da Universidade do Porto

Seeing with the sound:Sound-based human-context recognition using machine learning

Author: Wang Wei
Publication venue: University of Twente
Publication date: 13/10/2021
Field of study

University of Twente Research Information

Speech Recognition

Author
Publication venue: 'IntechOpen'
Publication date: 20/04/2021
Field of study

Chapters in the first part of the book cover all the essential speech processing techniques for building robust, automatic speech recognition systems: the representation for speech signals and the methods for speech-features extraction, acoustic and language modeling, efficient algorithms for searching the hypothesis space, and multimodal approaches to speech recognition. The last part of the book is devoted to other speech processing applications that can use the information from automatic speech recognition for speaker identification and tracking, for prosody modeling in emotion-detection systems and in other speech processing applications that are able to operate in real-world environments, like mobile communication services and smart homes

Directory of Open Access Books (DOAB)

Conference Proceedings of the Euroregio / BNAM 2022 Joint Acoustic Conference

Author
Publication venue: The European Acoustics Association (EAA)
Publication date: 09/05/2022
Field of study

VBN