Search CORE

3,601 research outputs found

Unstressed Vowels in German Learner English: An Instrumental Study

Author: Abercrombie
Abercrombie
Barry
Barry
Broselow
Broselow
Crystal
Crystal
Delattre
Delattre
Dellwo
Dellwo
Di
Di
Dretzke
Dretzke
Flege
Flege
German
German
German
German
Ghazali
Ghazali
Gibbon
Gibbon
Giegerich
Giegerich
Hansen Edwards
Hansen Edwards
Hoffmann
Hoffmann
Kaltenbacher
Kaltenbacher
Kendall
Kendall
Kohler
Kohler
Lee
Lee
Lukas Soenning
Mahwah
Mahwah
Nespor
Nespor
Ordin
Ordin
Parkes
Parkes
Pascoe
Pascoe
Peterson
Peterson
Pike
Pike
Porzuczek
Porzuczek
Prosodic
Prosodic
Roach
Roach
Traunmüller
Traunmüller
Wesener
Wesener
Wilcox
Wilcox
Publication venue: 'Walter de Gruyter GmbH'
Publication date: 01/01/2014
Field of study

This study investigates the production of vowels in unstressed syllables by advanced German learners of English in comparison with native speakers of Standard Southern British English. Two acoustic properties were measured: duration and formant structure. The results indicate that duration of unstressed vowels is similar in the two groups, though there is some variation depending on the phonetic context. In terms of formant structure, learners produce slightly higher F1 and considerably lower F2, the difference in F2 being statistically significant for each learner. Formant values varied as a function of context and orthographic representation of the vowel

Crossref

Biblioteka Nauki - repozytorium artykuÅÃ³w

Repozytorium Uniwersytetu Łódzkiego (University of Lodz Repository)

Language-independent talker-specificity in first-language and second-language speech production by bilingual talkers: L1 speaking rate predicts L2 speaking rate

Author: Blasingame Michael
Bradlow Ann R.
Kim Midam
Publication venue: 'Acoustical Society of America (ASA)'
Publication date: 08/11/2018
Field of study

Second-language (L2) speech is consistently slower than first-language (L1) speech, and L1 speaking rate varies within- and across-talkers depending on many individual, situational, linguistic, and sociolinguistic factors. It is asked whether speaking rate is also determined by a language-independent talker-specific trait such that, across a group of bilinguals, L1 speaking rate significantly predicts L2 speaking rate. Two measurements of speaking rate were automatically extracted from recordings of read and spontaneous speech by English monolinguals (n = 27) and bilinguals from ten L1 backgrounds (n = 86): speech rate (syllables/second), and articulation rate (syllables/second excluding silent pauses). Replicating prior work, L2 speaking rates were significantly slower than L1 speaking rates both across-groups (monolinguals' L1 English vs bilinguals' L2 English), and across L1 and L2 within bilinguals. Critically, within the bilingual group, L1 speaking rate significantly predicted L2 speaking rate, suggesting that a significant portion of inter-talker variation in L2 speech is derived from inter-talker variation in L1 speech, and that individual variability in L2 spoken language production may be best understood within the context of individual variability in L1 spoken language production

KU ScholarWorks

CAPT를 위한 발음 변이 분석 및 CycleGAN 기반 피드백 생성

Author: 양승희
Publication venue: 서울대학교 대학원
Publication date: 01/02/2020
Field of study

학위논문(박사)--서울대학교 대학원 :인문대학 협동과정 인지과학전공,2020. 2. 정민화.Despite the growing popularity in learning Korean as a foreign language and the rapid development in language learning applications, the existing computer-assisted pronunciation training (CAPT) systems in Korean do not utilize linguistic characteristics of non-native Korean speech. Pronunciation variations in non-native speech are far more diverse than those observed in native speech, which may pose a difficulty in combining such knowledge in an automatic system. Moreover, most of the existing methods rely on feature extraction results from signal processing, prosodic analysis, and natural language processing techniques. Such methods entail limitations since they necessarily depend on finding the right features for the task and the extraction accuracies. This thesis presents a new approach for corrective feedback generation in a CAPT system, in which pronunciation variation patterns and linguistic correlates with accentedness are analyzed and combined with a deep neural network approach, so that feature engineering efforts are minimized while maintaining the linguistically important factors for the corrective feedback generation task. Investigations on non-native Korean speech characteristics in contrast with those of native speakers, and their correlation with accentedness judgement show that both segmental and prosodic variations are important factors in a Korean CAPT system. The present thesis argues that the feedback generation task can be interpreted as a style transfer problem, and proposes to evaluate the idea using generative adversarial network. A corrective feedback generation model is trained on 65,100 read utterances by 217 non-native speakers of 27 mother tongue backgrounds. The features are automatically learnt in an unsupervised way in an auxiliary classifier CycleGAN setting, in which the generator learns to map a foreign accented speech to native speech distributions. In order to inject linguistic knowledge into the network, an auxiliary classifier is trained so that the feedback also identifies the linguistic error types that were defined in the first half of the thesis. The proposed approach generates a corrected version the speech using the learners own voice, outperforming the conventional Pitch-Synchronous Overlap-and-Add method.외국어로서의 한국어 교육에 대한 관심이 고조되어 한국어 학습자의 수가 크게 증가하고 있으며, 음성언어처리 기술을 적용한 컴퓨터 기반 발음 교육(Computer-Assisted Pronunciation Training; CAPT) 어플리케이션에 대한 연구 또한 적극적으로 이루어지고 있다. 그럼에도 불구하고 현존하는 한국어 말하기 교육 시스템은 외국인의 한국어에 대한 언어학적 특징을 충분히 활용하지 않고 있으며, 최신 언어처리 기술 또한 적용되지 않고 있는 실정이다. 가능한 원인으로써는 외국인 발화 한국어 현상에 대한 분석이 충분하게 이루어지지 않았다는 점, 그리고 관련 연구가 있어도 이를 자동화된 시스템에 반영하기에는 고도화된 연구가 필요하다는 점이 있다. 뿐만 아니라 CAPT 기술 전반적으로는 신호처리, 운율 분석, 자연어처리 기법과 같은 특징 추출에 의존하고 있어서 적합한 특징을 찾고 이를 정확하게 추출하는 데에 많은 시간과 노력이 필요한 실정이다. 이는 최신 딥러닝 기반 언어처리 기술을 활용함으로써 이 과정 또한 발전의 여지가 많다는 바를 시사한다. 따라서 본 연구는 먼저 CAPT 시스템 개발에 있어 발음 변이 양상과 언어학적 상관관계를 분석하였다. 외국인 화자들의 낭독체 변이 양상과 한국어 원어민 화자들의 낭독체 변이 양상을 대조하고 주요한 변이를 확인한 후, 상관관계 분석을 통하여 의사소통에 영향을 미치는 중요도를 파악하였다. 그 결과, 종성 삭제와 3중 대립의 혼동, 초분절 관련 오류가 발생할 경우 피드백 생성에 우선적으로 반영하는 것이 필요하다는 것이 확인되었다. 교정된 피드백을 자동으로 생성하는 것은 CAPT 시스템의 중요한 과제 중 하나이다. 본 연구는 이 과제가 발화의 스타일 변화의 문제로 해석이 가능하다고 보았으며, 생성적 적대 신경망 (Cycle-consistent Generative Adversarial Network; CycleGAN) 구조에서 모델링하는 것을 제안하였다. GAN 네트워크의 생성모델은 비원어민 발화의 분포와 원어민 발화 분포의 매핑을 학습하며, Cycle consistency 손실함수를 사용함으로써 발화간 전반적인 구조를 유지함과 동시에 과도한 교정을 방지하였다. 별도의 특징 추출 과정이 없이 필요한 특징들이 CycleGAN 프레임워크에서 무감독 방법으로 스스로 학습되는 방법으로, 언어 확장이 용이한 방법이다. 언어학적 분석에서 드러난 주요한 변이들 간의 우선순위는 Auxiliary Classifier CycleGAN 구조에서 모델링하는 것을 제안하였다. 이 방법은 기존의 CycleGAN에 지식을 접목시켜 피드백 음성을 생성함과 동시에 해당 피드백이 어떤 유형의 오류인지 분류하는 문제를 수행한다. 이는 도메인 지식이 교정 피드백 생성 단계까지 유지되고 통제가 가능하다는 장점이 있다는 데에 그 의의가 있다. 본 연구에서 제안한 방법을 평가하기 위해서 27개의 모국어를 갖는 217명의 유의미 어휘 발화 65,100개로 피드백 자동 생성 모델을 훈련하고, 개선 여부 및 정도에 대한 지각 평가를 수행하였다. 제안된 방법을 사용하였을 때 학습자 본인의 목소리를 유지한 채 교정된 발음으로 변환하는 것이 가능하며, 전통적인 방법인 음높이 동기식 중첩가산 (Pitch-Synchronous Overlap-and-Add) 알고리즘을 사용하는 방법에 비해 상대 개선률 16.67%이 확인되었다.Chapter 1. Introduction 1 1.1. Motivation 1 1.1.1. An Overview of CAPT Systems 3 1.1.2. Survey of existing Korean CAPT Systems 5 1.2. Problem Statement 7 1.3. Thesis Structure 7 Chapter 2. Pronunciation Analysis of Korean Produced by Chinese 9 2.1. Comparison between Korean and Chinese 11 2.1.1. Phonetic and Syllable Structure Comparisons 11 2.1.2. Phonological Comparisons 14 2.2. Related Works 16 2.3. Proposed Analysis Method 19 2.3.1. Corpus 19 2.3.2. Transcribers and Agreement Rates 22 2.4. Salient Pronunciation Variations 22 2.4.1. Segmental Variation Patterns 22 2.4.1.1. Discussions 25 2.4.2. Phonological Variation Patterns 26 2.4.1.2. Discussions 27 2.5. Summary 29 Chapter 3. Correlation Analysis of Pronunciation Variations and Human Evaluation 30 3.1. Related Works 31 3.1.1. Criteria used in L2 Speech 31 3.1.2. Criteria used in L2 Korean Speech 32 3.2. Proposed Human Evaluation Method 36 3.2.1. Reading Prompt Design 36 3.2.2. Evaluation Criteria Design 37 3.2.3. Raters and Agreement Rates 40 3.3. Linguistic Factors Affecting L2 Korean Accentedness 41 3.3.1. Pearsons Correlation Analysis 41 3.3.2. Discussions 42 3.3.3. Implications for Automatic Feedback Generation 44 3.4. Summary 45 Chapter 4. Corrective Feedback Generation for CAPT 46 4.1. Related Works 46 4.1.1. Prosody Transplantation 47 4.1.2. Recent Speech Conversion Methods 49 4.1.3. Evaluation of Corrective Feedback 50 4.2. Proposed Method: Corrective Feedback as a Style Transfer 51 4.2.1. Speech Analysis at Spectral Domain 53 4.2.2. Self-imitative Learning 55 4.2.3. An Analogy: CAPT System and GAN Architecture 57 4.3. Generative Adversarial Networks 59 4.3.1. Conditional GAN 61 4.3.2. CycleGAN 62 4.4. Experiment 63 4.4.1. Corpus 64 4.4.2. Baseline Implementation 65 4.4.3. Adversarial Training Implementation 65 4.4.4. Spectrogram-to-Spectrogram Training 66 4.5. Results and Evaluation 69 4.5.1. Spectrogram Generation Results 69 4.5.2. Perceptual Evaluation 70 4.5.3. Discussions 72 4.6. Summary 74 Chapter 5. Integration of Linguistic Knowledge in an Auxiliary Classifier CycleGAN for Feedback Generation 75 5.1. Linguistic Class Selection 75 5.2. Auxiliary Classifier CycleGAN Design 77 5.3. Experiment and Results 80 5.3.1. Corpus 80 5.3.2. Feature Annotations 81 5.3.3. Experiment Setup 81 5.3.4. Results 82 5.4. Summary 84 Chapter 6. Conclusion 86 6.1. Thesis Results 86 6.2. Thesis Contributions 88 6.3. Recommendations for Future Work 89 Bibliography 91 Appendix 107 Abstract in Korean 117 Acknowledgments 120Docto

SNU Open Repository and Archive

Recommended from our members

An exploratory study of foreign accent and phonological awareness in Korean learners of English

Author: Park Mi Sun
Publication venue: 'Columbia University Libraries/Information Services'
Publication date: 01/01/2019
Field of study

Communication in a second or multiple languages has become essential in the globalized world. However, acquiring a second language (L2) after a critical period is universally acknowledged to be challenging (Lenneberg, 1967). Late learners hardly reach a nativelike level in L2, particularly in its pronunciation, and their incomplete phonological acquisition is manifested by a foreign accent—a common and persistent feature of otherwise fluent L2 speech. Although foreign-accented speech is widespread, it has been a target of social constraints in L2-speaking communities, causing many learners and instructors to seek out ways to reduce foreign accents. Accordingly, research in L2 speech has unceasingly examined various learner-external and learner-internal factors of the occurrence of foreign accents as well as nonnative speech characteristics underlying the judgment of the degree of foreign accents. The current study aimed to expand the understanding of the characteristics and judgments of foreign accents by investigating phonological awareness, a construct pertinent to learners’ phonological knowledge, which has received little attention in research on foreign accents. The current study was exploratory and non-experimental research that targeted 40 adults with Korean-accented English living in the United States. The study first examined how 23 raters speaking American English as their native language detect, perceive, describe, and rate Korean-accented English. Through qualitative and quantitative analyses of the accent perception data, the study identified various phonological and phonetic deviations from the nativelike sounds, which largely result from the influence of first language (Korean) on L2 (English). The study then probed the relationship between foreign accents and learners’ awareness of the phonological system of L2, which was measured using production, perception, and verbalization tasks that tapped into the knowledge of L2 phonology. The study found a significant inverse relationship between the degree of a foreign accent and phonological awareness, particularly implicit knowledge of L2 segmentals. Further in-depth analyses revealed that explicit knowledge of L2 phonology alone was not sufficient for targetlike pronunciation. Findings suggest that L2 speakers experience varying degrees of difficulty in perceiving and producing different L2 segmentals, possibly resulting in foreign-accented speech

Columbia University Academic Commons

Measuring fluency: Temporal variables and pausing patterns in L2 English speech

Author: Park Soohwan
Publication venue: 'Purdue University (bepress)'
Publication date: 01/01/2016
Field of study

This paper examines temporal variables and pausing patterns in L2 English speech to investigate fluency as a measurable component of oral proficiency. Fluency can be defined as ‘speed and smoothness of oral delivery’. We can measure the speed of oral delivery through calculating temporal variables such as speech rate and mean syllables per run where ‘run’ is the vocal chunk between silent pauses. The smoothness of oral delivery can be measured through examination of pausing patterns by classifying the placement of pauses. Pauses may be placed in expected positions such as clause/phrase boundaries or in unexpected positions. Pause placement in unexpected positions may reduce the smoothness of oral delivery. The data sets are speech samples from the Oral English Proficiency Test (OEPT) but include the responses from two items (RAL: read aloud; NP: news passage). A total of 325 speakers across four different language groups (native speakers of Korean, Chinese, Hindi, and English) are represented across 6 proficiency levels (rated by holistic scoring based on the OEPT scale from 35 to 60). The speech samples were transcribed manually using a computer-assisted annotation tool that allowed capture of information about syllables, pausing boundaries, and types of pausing positions. Development of the annotation tool became a central concern of this study as establishing reliable and efficient methods in fluency research. Speech rate, mean syllables per run, and number of pauses per second were selected to examine temporal variables; number of unexpected pauses per second and expected pausing ratio were selected to compare pausing patterns across proficiency levels and language backgrounds. The results show that there are some linear relationships in temporal and pausing variables. High proficiency level speakers spoke at higher rates with expected pausing patterns compared to low proficiency level speakers who spoke at slower rates with almost no identifiable pausing patterns

Purdue E-Pubs

The development of automatic speech evaluation system for learners of English

Author: Kondo Yusuke
Publication venue
Publication date: 01/01/2010
Field of study

制度:新 ; 報告番号:甲3183号 ; 学位の種類:博士(教育学) ; 授与年月日:2010/11/30 ; 早大学位記番号:新547

Waseda University Repository

Recommended from our members

Heritage Languages: In the 'Wild' and in the Classroom

Author: Kagan Olga
Polinsky Maria
Publication venue: 'Wiley'
Publication date: 05/11/2009
Field of study

Heritage speakers are people raised in a home where one language is spoken who subsequently switch to another dominant language. The version of the home language that they have not completely acquired – heritage language – has only recently been given the attention it deserves from linguists and language instructors. Despite the appearance of great variation among heritage speakers, they fall along a continuum based upon the speakers' distance from the baseline language. Such a continuum-based model enables researchers and instructors to classify heritage speakers more accurately and readily. This article discusses the results of research on lower-proficiency speakers, identifying recurrent features of heritage languages in phonology, morphology, and syntax. Preliminary results indicate that different heritage languages share a number of structural similarities; this finding is important for the understanding of general processes involved in language acquisition. The article also presents implications of the main findings for language education and identifies areas needing further study.Linguistic

Harvard University - DASH

Recommended from our members

The role of HG in the analysis of temporal iteration and interaural correlation

Author: Barrett DJK
Hall DA
Publication venue
Publication date: 01/01/2004
Field of study

Nottingham Trent Institutional Repository (IRep)

Automatic Pronunciation Assessment -- A Review

Author: Ali Ahmed
Chowdhury Shammur Absar
Kheir Yassine El
Publication venue
Publication date: 21/10/2023
Field of study

Pronunciation assessment and its application in computer-aided pronunciation training (CAPT) have seen impressive progress in recent years. With the rapid growth in language processing and deep learning over the past few years, there is a need for an updated review. In this paper, we review methods employed in pronunciation assessment for both phonemic and prosodic. We categorize the main challenges observed in prominent research trends, and highlight existing limitations, and available resources. This is followed by a discussion of the remaining challenges and possible directions for future work.Comment: 9 pages, accepted to EMNLP Finding

arXiv.org e-Print Archive

Text reconstruction activities and teaching language forms

Author: Pawlak Mirosław
Publication venue: Wydawnictwo Uniwersytetu Łódzkiego
Publication date: 01/01/2011
Field of study

Even though there is a broad consensus that teaching language forms is facilitative or even necessary in some contexts, there are still disagreements concerning, among other things, how formal aspects of the target language should be taught. One important area of controversy is whether pedagogic intervention should be input-oriented, emphasizing comprehension of the form- meaning mappings represented by specific linguistic features or output-based, requiring learners to produce these features accurately in gradually more communicative activities. The present paper focuses on the latter of these two options and, basing on the claims of Swain‘s (1985, 1995) output hypothesis, it aims to demonstrates how text-reconstruction activities in which learners collaboratively produce written output trigger noticing, hypothesis-testing and metalinguistic reflection on language use. It presents a psycholinguistic and sociolinguistic rationale for the use of such tasks, discusses the types of such activities, provides an overview of research projects investigating their application and, finally, offers a set of implications for classroom use as well as suggestions for further research in this area

Repozytorium Uniwersytetu Łódzkiego (University of Lodz Repository)