Search CORE

44 research outputs found

Dyadic Speech-based Affect Recognition using DAMI-P2C Parent-child Multimodal Interaction Dataset

Author: Chung Junyoung
Fujita Yusuke
Gordon Goren
Hoover-Dempsey Kathleen V
Ioffe Sergey
Kamphaus Randy W
McNab Katrina
Neuman Susan B
Park Hae Won
Rudovic O.
Sainath Tara N
Spaulding Samuel
Stewart Angela
Trigeorgis George
Zhang Zixing
Zhao Huijuan
Publication venue: 'Association for Computing Machinery (ACM)'
Publication date: 20/08/2020
Field of study

Automatic speech-based affect recognition of individuals in dyadic conversation is a challenging task, in part because of its heavy reliance on manual pre-processing. Traditional approaches frequently require hand-crafted speech features and segmentation of speaker turns. In this work, we design end-to-end deep learning methods to recognize each person's affective expression in an audio stream with two speakers, automatically discovering features and time regions relevant to the target speaker's affect. We integrate a local attention mechanism into the end-to-end architecture and compare the performance of three attention implementations -- one mean pooling and two weighted pooling methods. Our results show that the proposed weighted-pooling attention solutions are able to learn to focus on the regions containing target speaker's affective information and successfully extract the individual's valence and arousal intensity. Here we introduce and use a "dyadic affect in multimodal interaction - parent to child" (DAMI-P2C) dataset collected in a study of 34 families, where a parent and a child (3-7 years old) engage in reading storybooks together. In contrast to existing public datasets for affect recognition, each instance for both speakers in the DAMI-P2C dataset is annotated for the perceived affect by three labelers. To encourage more research on the challenging task of multi-speaker affect sensing, we make the annotated DAMI-P2C dataset publicly available, including acoustic features of the dyads' raw audios, affect annotations, and a diverse set of developmental, social, and demographic profiles of each dyad.Comment: Accepted by the 2020 International Conference on Multimodal Interaction (ICMI'20

arXiv.org e-Print Archive

Crossref

BigSSL: Exploring the Frontier of Large-Scale Semi-Supervised Learning for Automatic Speech Recognition

Author: Beaufays Françoise
Cao Liangliang
Chan William
Chen Zhifeng
Chiu Chung-Cheng
Gulati Anmol
Han Wei
Huang Yanping
Jansen Aren
Le Quoc V.
Li Bo
Ma Min
Pang Ruoming
Park Daniel S.
Qin James
Ramabhadran Bhuvana
Sainath Tara N.
Shor Joel
Sim Khe Chai
Wang Shibo
Wang Yongqiang
Wu Yonghui
Xu Yuanzhong
Yu Jiahui
Zhang Yu
Zhou Zongwei
Publication venue: 'Institute of Electrical and Electronics Engineers (IEEE)'
Publication date: 21/07/2022
Field of study

We summarize the results of a host of efforts using giant automatic speech recognition (ASR) models pre-trained using large, diverse unlabeled datasets containing approximately a million hours of audio. We find that the combination of pre-training, self-training and scaling up model size greatly increases data efficiency, even for extremely large tasks with tens of thousands of hours of labeled data. In particular, on an ASR task with 34k hours of labeled data, by fine-tuning an 8 billion parameter pre-trained Conformer model we can match state-of-the-art (SoTA) performance with only 3% of the training data and significantly improve SoTA with the full training set. We also report on the universal benefits gained from using big pre-trained and self-trained models for a large set of downstream tasks that cover a wide range of speech domains and span multiple orders of magnitudes of dataset sizes, including obtaining SoTA performance on many public benchmarks. In addition, we utilize the learned representation of pre-trained networks to achieve SoTA results on non-ASR tasks.Comment: 14 pages, 7 figures, 13 tables; v2: minor corrections, reference baselines and bibliography updated; v3: corrections based on reviewer feedback, bibliography update

arXiv.org e-Print Archive

Antioxidant Status, Lipid Peroxidation and Testis-histoarchitecture Induced by Lead Nitrate and Mercury Chloride in Male Rats

Author: Acharya UR
Aebi H
Al-Attar AM
Apaydin FG
Apaydin FG
Apaydin FG
Bas H
Bas H
Demir F
Dirican EK
Etchevers A
Gharagozloo P
Habig WH
Hatice Bas
Janicka M
Jarup L
Kalender S
Kalender S
Kalender Y
Karaboduk H
Liu C
Marklund S
Mehrotra A
Messarah M
Ohkawa H
Orisakwe OE
Paglia DE
Pathak N
Pathak N
Patra RC
Plastunov B
Rainio MJ
Renugadevi J
Sainath SB
Sasaki JC
Sharma V
Suna Kalender
Uzun FG
Xia D
Yole M
Çelikoglu E
Publication venue: 'FapUNIFESP (SciELO)'
Publication date: 01/01/2016
Field of study

Crossref

Analisando audiências públicas no licenciamento ambiental: quem são e o que dizem os participantes sobre projetos de usinas de cana-de-açúcar

Crossref

A novel spatial-temporal prediction method for unsteady wake flows based on hybrid deep neural network

Author: Bouvrie J.
Goodfellow I.
Krizhevsky A.
Nair V.
Sainath T. N.
Shi X.
Zeiler M. D.
Zeiler M. D.
Publication venue: 'AIP Publishing'
Publication date
Field of study

Crossref

Anisotropic phenanthroline-based ruthenium polymers grafted on a titanium metal-organic framework for efficient photocatalytic hydrogen evolution

Author: Anjana Tripathi
Annadanam V. Sesha Sainath
Chandrani Nayak
Dibyendu Bhattacharyya
Gopinath Jonnalagadda
Ranjit Thapa
S. N. Jha
Saddam Sk
Spandana Gonuguntla
Ujjwal Pal
Vijayanand Perupogu
Publication venue: 'Springer Science and Business Media LLC'
Publication date: 01/12/2022
Field of study

Combining conjugated polymers with transition-metal-based metal-organic frameworks offers an opportunity to produce efficient photocatalytic materials. Here, exposed active sites and efficient charge transfer lead to hydrogen evolution rates of up to 2438 µmol g−1 h−1 for composites of anisotropic phenanthroline-based ruthenium polymers grafted on titanium-based MOFs

Directory of Open Access Journals