Search CORE

5 research outputs found

Gabor frames and deep scattering networks in audio processing

Author: Bammer Roswitha
Dörfler Monika
Harar Pavol
Publication venue: 'MDPI AG'
Publication date: 26/09/2019
Field of study

This paper introduces Gabor scattering, a feature extractor based on Gabor frames and Mallat's scattering transform. By using a simple signal model for audio signals specific properties of Gabor scattering are studied. It is shown that for each layer, specific invariances to certain signal characteristics occur. Furthermore, deformation stability of the coefficient vector generated by the feature extractor is derived by using a decoupling technique which exploits the contractivity of general scattering networks. Deformations are introduced as changes in spectral shape and frequency modulation. The theoretical results are illustrated by numerical examples and experiments. Numerical evidence is given by evaluation on a synthetic and a "real" data set, that the invariances encoded by the Gabor scattering transform lead to higher performance in comparison with just using Gabor transform, especially when few training samples are available.Comment: 26 pages, 8 figures, 4 tables. Repository for reproducibility: https://gitlab.com/hararticles/gs-gt . Keywords: machine learning; scattering transform; Gabor transform; deep learning; time-frequency analysis; CNN. Accepted and published after peer revisio

arXiv.org e-Print Archive

Multidisciplinary Digital Publishing Institute

Digital library of Brno University of Technology

Basic Filters for Convolutional Neural Networks Applied to Music: Training or Design?

Author: Doerfler Monika
Grill Thomas
Bammer Roswitha
Flexer Arthur
Publication venue
Publication date: 05/07/2017
Field of study

When convolutional neural networks are used to tackle learning problems based on music or, more generally, time series data, raw one-dimensional data are commonly pre-processed to obtain spectrogram or mel-spectrogram coefficients, which are then used as input to the actual neural network. In this contribution, we investigate, both theoretically and experimentally, the influence of this pre-processing step on the network's performance and pose the question, whether replacing it by applying adaptive or learned filters directly to the raw data, can improve learning success. The theoretical results show that approximately reproducing mel-spectrogram coefficients by applying adaptive filters and subsequent time-averaging is in principle possible. We also conducted extensive experimental work on the task of singing voice detection in music. The results of these experiments show that for classification based on Convolutional Neural Networks the features obtained from adaptive filter banks followed by time-averaging perform better than the canonical Fourier-transform-based mel-spectrogram coefficients. Alternative adaptive approaches with center frequencies or time-averaging lengths learned from training data perform equally well.Comment: Completely revised version; 21 pages, 4 figure

arXiv.org e-Print Archive

Dryad Digital Repository (Duke University)

Gabor Frames and Deep Scattering Networks in Audio Processing

Author: Bammer Roswitha (Faculty of Mathematics, Faculty of Mathematics, University of Vienna)
Dörfler Monika (Faculty of Mathematics, Faculty of Mathematics, University of Vienna)
Harar Pavol (Faculty of Mathematics, Faculty of Mathematics, University of Vienna)
Publication venue: 'MDPI AG'
Publication date: 01/01/2019
Field of study

This paper introduces Gabor scattering, a feature extractor based on Gabor frames and Mallat’s scattering transform. By using a simple signal model for audio signals, specific properties of Gabor scattering are studied. It is shown that, for each layer, specific invariances to certain signal characteristics occur. Furthermore, deformation stability of the coefficient vector generated by the feature extractor is derived by using a decoupling technique which exploits the contractivity of general scattering networks. Deformations are introduced as changes in spectral shape and frequency modulation. The theoretical results are illustrated by numerical examples and experiments. Numerical evidence is given by evaluation on a synthetic and a “real” dataset, that the invariances encoded by the Gabor scattering transform lead to higher performance in comparison with just using Gabor transform, especially when few training samples are available.© 2019 by the author

Permanent Hosting, Archiving and Indexing of Digital Resources and Assets