10,604 research outputs found
Large-scale Hierarchical Alignment for Data-driven Text Rewriting
We propose a simple unsupervised method for extracting pseudo-parallel
monolingual sentence pairs from comparable corpora representative of two
different text styles, such as news articles and scientific papers. Our
approach does not require a seed parallel corpus, but instead relies solely on
hierarchical search over pre-trained embeddings of documents and sentences. We
demonstrate the effectiveness of our method through automatic and extrinsic
evaluation on text simplification from the normal to the Simple Wikipedia. We
show that pseudo-parallel sentences extracted with our method not only
supplement existing parallel data, but can even lead to competitive performance
on their own.Comment: RANLP 201
Detecting event-related recurrences by symbolic analysis: Applications to human language processing
Quasistationarity is ubiquitous in complex dynamical systems. In brain
dynamics there is ample evidence that event-related potentials reflect such
quasistationary states. In order to detect them from time series, several
segmentation techniques have been proposed. In this study we elaborate a recent
approach for detecting quasistationary states as recurrence domains by means of
recurrence analysis and subsequent symbolisation methods. As a result,
recurrence domains are obtained as partition cells that can be further aligned
and unified for different realisations. We address two pertinent problems of
contemporary recurrence analysis and present possible solutions for them.Comment: 24 pages, 6 figures. Draft version to appear in Proc Royal Soc
- …