Search CORE

41,982 research outputs found

Finding the Most Uniform Changes in Vowel Polygon Caused by Psychological Stress

Author: Sigmund M.
Stanek M.
Publication venue: 'Brno University of Technology'
Publication date: 01/06/2015
Field of study

Using vowel polygons, exactly their parameters, is chosen as the criterion for achievement of differences between normal state of speaker and relevant speech under real psychological stress. All results were experimentally obtained by created software for vowel polygon analysis applied on ExamStress database. Selected 6 methods based on cross-correlation of different features were classified by the coefficient of variation and for each individual vowel polygon, the efficiency coefficient marking the most significant and uniform differences between stressed and normal speech were calculated. As the best method for observing generated differences resulted method considered mean of cross correlation values received for difference area value with vector length and angle parameter couples. Generally, best results for stress detection are achieved by vowel triangles created by /i/-/o/-/u/ and /a/-/i/-/o/ vowel triangles in formant planes containing the fifth formant F5 combined with other formants

Directory of Open Access Journals

Digital library of Brno University of Technology

Estimation of glottal closure instants in voiced speech using the DYPSA algorithm

Author: Brookes M
Gudnason J
Kounoudes A
Naylor PA
Publication venue: 'Institute of Electrical and Electronics Engineers (IEEE)'
Publication date: 01/01/2007
Field of study

Published versio

Spiral - Imperial College Digital Repository

Speaker-independent emotion recognition exploiting a psychologically-inspired binary cascade classification schema

Author: A. Austermann
B. Schuller
B. Schuller
B. Schuller
B. Schuller
B. Schuller
B. Schuller
B. Yang
C. Busso
C. M. Lee
C. Nass
D. Bitouk
D. J. C. MacKay
D. Ververidis
D. Ververidis
D. Watson
E. Benetos
E. Benetos
E. Fersini
E. I. Konstantinidis
F. Burkhardt
F. Burkhardt
Fabio Paternò
H. Altun
H. Gunes
H. K. Mishra
H. Mixdorff
H. P. Espinosa
I. Guyon
I. Guyon
I. R. Murray
J. D. Markel
J. Hirschberg
J. Pittermann
K. Dai
K. R. Scherer
L. B. Jackson
M. Ayadi El
M. Kotti
M. Kotti
M. M. Sondhi
M. Pantic
M. Pantic
Margarita Kotti
N. Sato
N. Vanello
P. Boersma
P. Ekman
P. Ekman
P. N. Juslin
P. Ruvolo
P. Zervas
R. A. Calvo
R. Cowie
R. Tato
R. W. Picard
S. Chandaka
S. Ntalampiras
T. Iliou
T. L. Pao
T. P. Kostoulas
T. Vogt
W. Bosma
W. Minker
Z. Inanoglu
Z. Zeng
Publication venue: 'Springer Science and Business Media LLC'
Publication date: 01/06/2012
Field of study

In this paper, a psychologically-inspired binary cascade classification schema is proposed for speech emotion recognition. Performance is enhanced because commonly confused pairs of emotions are distinguishable from one another. Extracted features are related to statistics of pitch, formants, and energy contours, as well as spectrum, cepstrum, perceptual and temporal features, autocorrelation, MPEG-7 descriptors, Fujisakis model parameters, voice quality, jitter, and shimmer. Selected features are fed as input to K nearest neighborhood classifier and to support vector machines. Two kernels are tested for the latter: Linear and Gaussian radial basis function. The recently proposed speaker-independent experimental protocol is tested on the Berlin emotional speech database for each gender separately. The best emotion recognition accuracy, achieved by support vector machines with linear kernel, equals 87.7%, outperforming state-of-the-art approaches. Statistical analysis is first carried out with respect to the classifiers error rates and then to evaluate the information expressed by the classifiers confusion matrices. © Springer Science+Business Media, LLC 2011

Crossref

Spiral - Imperial College Digital Repository

The acquisition of English L2 prosody by Italian native speakers: experimental data and pedagogical implications

Author: Busa' Maria Grazia
Stella A.
Publication venue: University of Santa Barbara
Publication date: 01/01/2015
Field of study

This paper investigates Yes-No question intonation patterns in English L2, Italian L1, and English L1. The aim is to test the hypothesis that L2 learners may show different acquisition strategies for different dimensions of intonation, and particularly the phonological and phonetic components. The study analyses the nuclear intonation contours of 4 target English words and 4 comparable Italian words consisting of sonorant segments, stressed on the semi-final or final syllable, and occurring in Yes-No questions in sentence-final position (e.g., Will you attend the memorial?, Hai sentito la Melania?). The words were contained in mini-dialogues of question-answer pairs, and read 5 times by 4 Italian speakers (Padova area, North-East Italy) and 3 English female speakers (London area, UK). The results show that: 1) different intonation patterns may be used to realize the same grammatical function; 2) different developmental processes are at work, including transfer of L1 categories and the acquisition of L2 phonological categories. These results suggest that the phonetic dimension of L2 intonation may be more difficult to learn than the phonological one

Archivio istituzionale della ricerca - Università di Padova

How do you say ‘hello’? Personality impressions from brief novel voices

Author: A Todorov
A Todorov
AC Little
AC Little
AC Little
Alexander Todorov
AW Young
B Yao
B Yao
BC Jones
C Ferdenzi
C Nass
CA Klofstad
CA Sutherland
CC Tigue
CD Aronovitch
Charles R. Larson
CL Apicella
CL Apicella
CP Said
CT Ferrand
CY Olivola
D Rendall
DA Kenny
DA Leopold
DA Puts
DC Funder
DR Feinberg
DR Feinberg
DS Berry
DS Berry
DS Berry
DS Berry
E Vannoni
FT Passini
G Rhodes
G Rhodes
GW Allport
IR Titze
IS Penton-Voak
J Kreiman
J Willis
JH Langlois
JH Langlois
JJ Horton
JJ Ohala
JM Montepare
JS Wiggins
JW Lewis
K Grammer
K Miyake
KR Scherer
L Bruckert
L Bruckert
L Germine
LA Zebrowitz
LA Zebrowitz
LA Zebrowitz
LA Zebrowitz
LA Zebrowitz-McArthur
LZ McArthur
M Latinus
M Latinus
M Latinus
M Shevlin
M Zuckerman
M Zuckerman
M Zuckerman
NJ Lass
NJ Lass
NN Oosterhof
O Baumann
P Belin
P Belin
P Boersma
P Borkenau
Pascal Belin
PE Bestelmeyer
PEG Bestelmeyer
Phil McAleer
R Hassin
RA Page
RM Krauss
RR Mccrae
RS Kramer
S Evans
S Patel
S Rosenberg
SA Collins
SC Verosky
SJ Ko
SM Hughes
SM Hughes
SM Hughes
ST Fiske
TK Perrachione
V Bruce
V Pivonkova
VX Luevano
WA van Dommelen
WA van Dommelen
WT Fitch
WT Fitch
WT Norman
Publication venue: 'Public Library of Science (PLoS)'
Publication date: 01/01/2014
Field of study

On hearing a novel voice, listeners readily form personality impressions of that speaker. Accurate or not, these impressions are known to affect subsequent interactions; yet the underlying psychological and acoustical bases remain poorly understood. Furthermore, hitherto studies have focussed on extended speech as opposed to analysing the instantaneous impressions we obtain from first experience. In this paper, through a mass online rating experiment, 320 participants rated 64 sub-second vocal utterances of the word ‘hello’ on one of 10 personality traits. We show that: (1) personality judgements of brief utterances from unfamiliar speakers are consistent across listeners; (2) a two-dimensional ‘social voice space’ with axes mapping Valence (Trust, Likeability) and Dominance, each driven by differing combinations of vocal acoustics, adequately summarises ratings in both male and female voices; and (3) a positive combination of Valence and Dominance results in increased perceived male vocal Attractiveness, whereas perceived female vocal Attractiveness is largely controlled by increasing Valence. Results are discussed in relation to the rapid evaluation of personality and, in turn, the intent of others, as being driven by survival mechanisms via approach or avoidance behaviours. These findings provide empirical bases for predicting personality impressions from acoustical analyses of short utterances and for generating desired personality impressions in artificial voices

Crossref

HAL AMU

Directory of Open Access Journals

PubMed Central

Enlighten

FigShare

Detecting User Engagement in Everyday Conversations

Author: Aoki Paul M.
Woodruff Allison
Yu Chen
Publication venue
Publication date: 01/01/2004
Field of study

This paper presents a novel application of speech emotion recognition: estimation of the level of conversational engagement between users of a voice communication system. We begin by using machine learning techniques, such as the support vector machine (SVM), to classify users' emotions as expressed in individual utterances. However, this alone fails to model the temporal and interactive aspects of conversational engagement. We therefore propose the use of a multilevel structure based on coupled hidden Markov models (HMM) to estimate engagement levels in continuous natural speech. The first level is comprised of SVM-based classifiers that recognize emotional states, which could be (e.g.) discrete emotion types or arousal/valence levels. A high-level HMM then uses these emotional states as input, estimating users' engagement in conversation by decoding the internal states of the HMM. We report experimental results obtained by applying our algorithms to the LDC Emotional Prosody and CallFriend speech corpora.Comment: 4 pages (A4), 1 figure (EPS

arXiv.org e-Print Archive

CiteSeerX