Search CORE

604 research outputs found

Cue Phrase Classification Using Machine Learning

Author: Litman Diane J.
Publication venue
Publication date: 01/01/1996
Field of study

Cue phrases may be used in a discourse sense to explicitly signal discourse structure, but also in a sentential sense to convey semantic rather than structural information. Correctly classifying cue phrases as discourse or sentential is critical in natural language processing systems that exploit discourse structure, e.g., for performing tasks such as anaphora resolution and plan recognition. This paper explores the use of machine learning for classifying cue phrases as discourse or sentential. Two machine learning programs (Cgrendel and C4.5) are used to induce classification models from sets of pre-classified cue phrases and their features in text and speech. Machine learning is shown to be an effective technique for not only automating the generation of classification models, but also for improving upon previous results. When compared to manually derived classification models already in the literature, the learned models often perform with higher accuracy and contain new linguistic insights into the data. In addition, the ability to automatically construct classification models makes it easier to comparatively analyze the utility of alternative feature representations of the data. Finally, the ease of retraining makes the learning approach more scalable and flexible than manual methods.Comment: 42 pages, uses jair.sty, theapa.bst, theapa.st

arXiv.org e-Print Archive

CiteSeerX

An infrastructure for Turkish prosody generation in text-to-speech synthesis

Author: Kulekci M. Oguzhan
Külekçi M. Oğuzhan
Oflazer Kemal
Publication venue
Publication date: 01/06/2006
Field of study

Text-to-speech engines benefit from natural language processing while generating the appropriate prosody. In this study, we investigate the natural language processing infrastructure for Turkish prosody generation in three steps as pronunciation disambiguation, phonological phrase detection and intonation level assignment. We focus on phrase boundary detection and intonation assignment. We propose a phonological phrase detection scheme based on syntactic analysis for Turkish and assign one of three intonation levels to words in detected phrases. Empirical observations on 100 sentences show that the proposed scheme works with approximately 85% accuracy

Sabanci University Research Database

Recommended from our members

Empirical Studies on the Disambiguation of Cue Phrases

Author: Hirschberg Julia Bell
Litman Diane
Publication venue: 'Columbia University Libraries/Information Services'
Publication date: 01/01/1991
Field of study

Cue phrases are linguistic expressions such as now and well that function as explicit indicators of the structure of a discourse. For example, now may signal the beginning of a subtopic or a return to a previous topic, while well may mark subsequent material as a response to prior material, or as an explanatory comment. However, while cue phrases may convey discourse structure, each also has one or more alternate uses. While incidentally may be used sententially as an adverbial, for example, the discourse use initiates a digression. Although distinguishing discourse and sentential uses of cue phrases is critical to the interpretation and generation of discourse, the question of how speakers and hearers accomplish this disambiguation is rarely addressed. This paper reports results of empirical studies on discourse and sentential uses of cue phrases, in which both text-based and prosodic features were examined for disambiguating power. Based on these studies, it is proposed that discourse versus sentential usage may be distinguished by intonational features, specifically, pitch accent and prosodic phrasing. A prosodic model that characterizes these distinctions is identified. This model is associated with features identifiable from text analysis, including orthography and part of speech, to permit the application of the results of the prosodic analysis to the generation of appropriate intonational features for discourse and sentential uses of cue phrases in synthetic speech

Columbia University Academic Commons

Exploring complex vowels as phrase break correlates in a corpus of English speech with ProPOSEL, a prosody and POS English lexicon

Author: Atwell E
Brierley C
Publication venue
Publication date: 01/01/2009
Field of study

Real-world knowledge of syntax is seen as integral to the machine learning task of phrase break prediction but there is a deficiency of a priori knowledge of prosody in both rule-based and data-driven classifiers. Speech recognition has established that pauses affect vowel duration in preceding words. Based on the observation that complex vowels occur at rhythmic junctures in poetry, we run significance tests on a sample of transcribed, contemporary British English speech and find a statistically significant correlation between complex vowels and phrase breaks. The experiment depends on automatic text annotation via ProPOSEL, a prosody and part-of-speech English lexicon. Copyright © 2009 ISCA

White Rose Research Online

Leeds Beckett Repository

Secondary stress in Brazilian Portuguese: the interplay between production and perception studies

Author: Arantes Pablo
Barbosa Plinio A.
Publication venue: Technische Universität Dresden Press
Publication date: 01/01/2006
Field of study

This paper reports experiments on speech production showing that secondary stress in Brazilian Portuguese (BP) can be best described as phrase-initial prominence cued by greater duration and pitch accent excursion in initial position. It also reports a perception experiment in which clicks were associated to consecutive V-to-V positions in stress groups. Mean click detection RTs are gradient, but show no influence of initial lengthening. RTs near the phrasally stressed position are shorter and almost 60% of RT variance can be accounted for by produced timing patterns

CogPrints Cognitive Sciences Eprint Archive

Pauses and the temporal structure of speech

Author: Zellner Brigitte
Publication venue: John Wiley
Publication date: 01/01/1994
Field of study

Natural-sounding speech synthesis requires close control over the temporal structure of the speech flow. This includes a full predictive scheme for the durational structure and in particuliar the prolongation of final syllables of lexemes as well as for the pausal structure in the utterance. In this chapter, a description of the temporal structure and the summary of the numerous factors that modify it are presented. In the second part, predictive schemes for the temporal structure of speech ("performance structures") are introduced, and their potential for characterising the overall prosodic structure of speech is demonstrated

CiteSeerX

CogPrints Cognitive Sciences Eprint Archive

Unsupervised continuous-valued word features for phrase-break prediction without a part-of-speech tagger.

Author: King Simon
Watts Oliver
Yamagishi Junichi
Publication venue
Publication date: 01/08/2011
Field of study

Edinburgh Research Explorer

Tagging Prosody and Discourse Structure in Elicited Spontaneous Speech

Author: Beckman Mary E.
Venditti Jennifer J.
Publication venue: Ohio State University. Department of Linguistics
Publication date: 01/01/2000
Field of study

This paper motivates and describes the annotation and analysis of prosody and discourse structure for several large spoken language corpora. The annotation schema are of two types: tags for prosody and intonation, and tags for several aspects of discourse structure. The choice of the particular tagging schema in each domain is based in large part on the insights they provide in corpus-based studies of the relationship between discourse structure and the accenting of referring expressions in American English. We first describe these results and show that the same models account for the accenting of pronouns in an extended passage from one of the Speech Warehouse hotel-booking dialogues. We then turn to corpora described in Venditti [Ven00], which adapts the same models to Tokyo Japanese. Japanese is interesting to compare to English, because accent is lexically specified and so cannot mark discourse focus in the same way. Analyses of these corpora show that local pitch range expansion serves the analogous focusing function in Japanese. The paper concludes with a section describing several outstanding questions in the annotation of Japanese intonation which corpus studies can help to resolve.Work reported in this paper was supported in part by a grant from the Ohio State University Office of Research, to Mary E. Beckman and co-principal investigators on the OSU Speech Warehouse project, and by an Ohio State University Presidential Fellowship to Jennifer J. Venditti

KnowledgeBank at OSU