Search CORE

437 research outputs found

Synapse at CAp 2017 NER challenge: Fasttext CRF

Author: Alexandra J. Weisberg (4234153)
Briana S. Bullington (4234156)
Eric R. Moore (4234150)
Jeff Chang (228277)
Kimberly H. Halsey (208487)
Yuan Jiang (296541)
Publication venue
Publication date: 01/01/2017
Field of study

We present our system for the CAp 2017 NER challenge which is about named entity recognition on French tweets. Our system leverages unsupervised learning on a larger dataset of French tweets to learn features feeding a CRF model. It was ranked first without using any gazetteer or structured external data, with an F-measure of 58.89\%. To the best of our knowledge, it is the first system to use fasttext embeddings (which include subword representations) and an embedding-based sentence representation for NER

arXiv.org e-Print Archive

Directory of Open Access Journals

FigShare

Towards the ontology-based approach for factual information matching

Author: Cherednichenko Olga
Doroshenko Anastsiia
Sharonova Natalia Valeriyevna
Publication venue: Друкарня Мадрид
Publication date: 01/01/2018
Field of study

Factual information is information based on facts or relating to facts. The reliability of automatically extracted facts is the main problem of processing factual information. The fact retrieval system remains one of the most effective tools for identifying the information for decision-making. In this work, we explore how can natural language processing methods and problem domain ontology help to check contradictions and mismatches in facts automatically

Electronic National Technical University "Kharkiv Polytechnic Institute" Institutional Repository (eNTUKhPIIR)

On the Need for Structure Modelling in Sequence Prediction

Author: Diethe Tom
Flach Peter
Twomey Niall
Publication venue: 'Springer Science and Business Media LLC'
Publication date: 21/07/2016
Field of study

Crossref

Springer - Publisher Connector

Explore Bristol Research

An automatic part-of-speech tagger for Middle Low German

Author: Breitbarth Anne
Desmet Bart
Farasyn Melissa
Hoste Veronique
Koleva Mariya
Publication venue: 'John Benjamins Publishing Company'
Publication date: 01/01/2017
Field of study

Syntactically annotated corpora are highly important for enabling large-scale diachronic and diatopic language research. Such corpora have recently been developed for a variety of historical languages, or are still under development. One of those under development is the fully tagged and parsed Corpus of Historical Low German (CHLG), which is aimed at facilitating research into the highly under-researched diachronic syntax of Low German. The present paper reports on a crucial step in creating the corpus, viz. the creation of a part-of-speech tagger for Middle Low German (MLG). Having been transmitted in several non-standardised written varieties, MLG poses a challenge to standard POS taggers, which usually rely on normalized spelling. We outline the major issues faced in the creation of the tagger and present our solutions to them

Crossref

Ghent University Academic Bibliography

Dutch named entity recognition using ensemble classifiers

Author: Desmet Bart
Hoste Veronique
Publication venue: Landelijke Onderzoeksschool Taalwetenschap (LOT)
Publication date: 01/01/2010
Field of study

Ghent University Academic Bibliography