Search CORE

716 research outputs found

MIRACLE Retrieval Experiments with East Asian Languages

Author: González Cristóbal José Carlos
Goñi Menoyo José Miguel
Martínez Fernández José Luis
Villena Román Julio
Publication venue: E.T.S.I. Telecomunicación (UPM)
Publication date: 01/01/2005
Field of study

This paper describes the participation of MIRACLE in NTCIR 2005 CLIR task. Although our group has a strong background and long expertise in Computational Linguistics and Information Retrieval applied to European languages and using Latin and Cyrillic alphabets, this was our first attempt on East Asian languages. Our main goal was to study the particularities and distinctive characteristics of Japanese, Chinese and Korean, specially focusing on the similarities and differences with European languages, and carry out research on CLIR tasks which include those languages. The basic idea behind our participation in NTCIR is to test if the same familiar linguisticbased techniques may also applicable to East Asian languages, and study the necessary adaptations

Archivo Digital UPM

LCC-DCU C-C question answering task at NTCIR-5

Author: Jones Gareth J.F.
Wang Bin
Publication venue
Publication date: 01/12/2005
Field of study

This paper describes the work for our participation in the NTCIR-5 Chinese to Chinese Question Answering task. Our strategy is based on the “Retrieval plus Extraction” approach. We first retrieve relevant documents, then retrieve short passages from the above documents, and finally extract named entity answers from the most relevant passages. For question type identification, we use simple heuristic rules which can cover most questions. The Lemur toolkit with the OKAPI model is used for document retrieval. Results of our task submission are given and some preliminary conclusions drawn

Irish Universities

DCU Online Research Access Service

An analysis of question processing of English and Chinese for the NTCIR 5 cross-language question answering task

Author: Guo Yuqing
Jones Gareth J.F.
Judge John
Wang Bin
Publication venue: 'National Institute of Informatics (NII)'
Publication date: 01/01/2005
Field of study

An important element in question answering systems is the analysis and interpretation of questions. Using the NTCIR 5 Cross-Language Question Answering (CLQA) question test set we demonstrate that the accuracy of deep question analysis is dependent on the quantity and suitability of the available linguistic resources. We further demonstrate that applying question analysis tools developed on monolingual training materials to questions translated Chinese-English and English-Chinese using machine translation produces much reduced effectiveness in interpretation of the question. This latter result indicates that question analysis for CLQA should primarily be conducted in the question language prior to translation

CiteSeerX

Irish Universities

DCU Online Research Access Service

CRL at Ntcir2

Author: Isahara Hitoshi
Ma Qing
Murata Masaki
Ozaku Hiromi
Utiyama Masao
Publication venue
Publication date: 01/01/2001
Field of study

We have developed systems of two types for NTCIR2. One is an enhenced version of the system we developed for NTCIR1 and IREX. It submitted retrieval results for JJ and CC tasks. A variety of parameters were tried with the system. It used such characteristics of newspapers as locational information in the CC tasks. The system got good results for both of the tasks. The other system is a portable system which avoids free parameters as much as possible. The system submitted retrieval results for JJ, JE, EE, EJ, and CC tasks. The system automatically determined the number of top documents and the weight of the original query used in automatic-feedback retrieval. It also determined relevant terms quite robustly. For EJ and JE tasks, it used document expansion to augment the initial queries. It achieved good results, except on the CC tasks.Comment: 11 pages. Computation and Language. This paper describes our results of information retrieval in the NTCIR2 contes

arXiv.org e-Print Archive

CiteSeerX

Language Modeling for Multi-Domain Speech-Driven Text Retrieval

Author: Fujii Atsushi
Ishikawa Tetsuya
Itou Katunobu
Publication venue: 'Institute of Electrical and Electronics Engineers (IEEE)'
Publication date: 01/01/2001
Field of study

We report experimental results associated with speech-driven text retrieval, which facilitates retrieving information in multiple domains with spoken queries. Since users speak contents related to a target collection, we produce language models used for speech recognition based on the target collection, so as to improve both the recognition and retrieval accuracy. Experiments using existing test collections combined with dictated queries showed the effectiveness of our method

arXiv.org e-Print Archive

CiteSeerX

Crossref

A Survey on Retrieval of Mathematical Knowledge

Author: A Asperti
A Asperti
A Kohlhase
A Kohlhase
AM Youssef
AS Youssef
AS Youssef
BR Miller
BR Miller
BR Miller
D Delahaye
F Guidi
F Rabe
G Bancerek
G Bancerek
G Bancerek
I Normann
M Adeel
M Líška
M-Q Nghiem
ME Altamimi
O Caprotti
P Baumgartner
P Cairns
P Libbrecht
P Libbrecht
P Libbrecht
Q Zhang
R Miner
R Zanibbi
S Kamali
T Gauthier
Y Haralambous
Publication venue
Publication date: 01/01/2015
Field of study

We present a short survey of the literature on indexing and retrieval of mathematical knowledge, with pointers to 72 papers and tentative taxonomies of both retrieval problems and recurring techniques.Comment: CICM 2015, 20 page

arXiv.org e-Print Archive

Crossref

Archivio istituzionale della ricerca - Alma Mater Studiorum Università di Bologna

ICT-DCU question answering task at NTCIR-6

Author: Jones Gareth J.F.
Wang Bin
Zhang Sen
Publication venue
Publication date: 01/01/2007
Field of study

This paper describes details of our participation in the NTCIR-6 Chinese-to-Chinese Question Answering task. We use the “retrieval plus extraction approach” to get answers for questions. We first split the documents into short passages, and then retrieve potentially relevant passages for a question, and finally extract named entity answers from the most relevant passages. For question type identification, we use simple heuristic rules which cover most questions. The Lemur toolkit was used with the okapi model for document retrieval. Results of our task submission are given and some preliminary conclusions drawn

CiteSeerX

Irish Universities

DCU Online Research Access Service

Which one is better: presentation-based or content-based math search?

Author: A.S. Youssef
B.R. Miller
J. Mišutka
M. Adeel
M. Kohlhase
M.E. Altamimi
M.Q. Nghiem
R. Miner
R. Zanibbi
S. Kamali
Publication venue
Publication date: 01/01/2014
Field of study

Mathematical content is a valuable information source and retrieving this content has become an important issue. This paper compares two searching strategies for math expressions: presentation-based and content-based approaches. Presentation-based search uses state-of-the-art math search system while content-based search uses semantic enrichment of math expressions to convert math expressions into their content forms and searching is done using these content-based expressions. By considering the meaning of math expressions, the quality of search system is improved over presentation-based systems

arXiv.org e-Print Archive

CiteSeerX

Crossref