Search CORE

1 research outputs found

Comparative Study on Sentence Boundary Prediction for German and English Broadcast News

Author: Garner Philip N.
Imseng David
Lazaridis Alexandros
Nanchen Alexandre
Wang Yang
Publication venue: Idiap
Publication date: 19/07/2017
Field of study

We present a comparative study on sentence boundary prediction for German and English broadcast news that explores generalization across different languages. In the feature extraction stage, word pause duration is firstly extracted from word aligned speech, and forward and backward language models are utilized to extract textual features. Then a gradient boosted machine is optimized by grid search to map these features to punctuation marks. Experimental results confirm that word pause duration is a simple yet effective feature to predict whether there is a sentence boundary after that word. We found that Bayes risk derived from pause duration distributions of sentence boundary words and non-boundary words is an effective measure to assess the inherent difficulty of sentence boundary prediction. The proposed method achieved F-measures of over 90% on reference text and around 90% on ASR transcript for both German broadcast news corpus and English multi-genre broadcast news corpus. This demonstrates the state of the art performance of the proposed method

Infoscience - École polytechnique fédérale de Lausanne