Search CORE

182 research outputs found

DCU-TCD@LogCLEF 2010: re-ranking document collections and query performance estimation

Author: Ghorab M. Rami
Jones Gareth J.F.
Leveling Johannes
Magdy Walid
Wade Vincent
Publication venue
Publication date: 01/09/2010
Field of study

This paper describes the collaborative participation of Dublin City University and Trinity College Dublin in LogCLEF 2010. Two sets of experiments were conducted. First, different aspects of the TEL query logs were analysed after extracting user sessions of consecutive queries on a topic. The relation between the queries and their length (number of terms) and position (first query or further reformulations) was examined in a session with respect to query performance estimators such as query scope, IDF-based measures, simplified query clarity score, and average inverse document collection frequency. Results of this analysis suggest that only some estimator values show a correlation with query length or position in the TEL logs (e.g. similarity score between collection and query). Second, the relation between three attributes was investigated: the user's country (detected from IP address), the query language, and the interface language. The investigation aimed to explore the influence of the three attributes on the user's collection selection. Moreover, the investigation involved assigning different weights to the three attributes in a scoring function that was used to re-rank the collections displayed to the user according to the language and country. The results of the collection re-ranking show a significant improvement in Mean Average Precision (MAP) over the original collection ranking of TEL. The results also indicate that the query language and interface language have more in uence than the user's country on the collections selected by the users

CiteSeerX

Irish Universities

DCU Online Research Access Service

IdSay: question answering for portuguese

Author: Carvalho Gracinda
Matos David Martins de
Rocio Vitor
Publication venue: 'Springer Science and Business Media LLC'
Publication date: 01/01/2009
Field of study

IdSay is an open domain Question Answering (QA) system for Portuguese. Its current version can be considered a baseline version, using mainly techniques from the area of Information Retrieval (IR). The only external information it uses besides the text collections is lexical information for Portuguese. It was submitted to the monolingual Portuguese task of the QA track of the Cross-Language Evaluation Forum 2008 (QA@CLEF) for the first time, and it answered correctly to 65 of the 200 questions in the first answer, and to 85 answers considering the three answers that could be returned per question. Generally, the types of questions that are answered better by IdSay system are measure factoids, count factoids and definitions, but there is still work to be done in these areas, as well as in the treatment of time. List questions, location and people/organization factoids are the types of question with more room for improvement.info:eu-repo/semantics/publishedVersio

Crossref

Repositório Aberto da Universidade Aberta

Combining Textual and Visual Information for Image Retrieval in the Medical Domain

Author: Anna Morou
Chatzichristofis SA
Jones KS
Tamura H
Theodore Kalamboukis
Vogt CC
Wu S
Wu S
Wu S
Yiannis Gkoufas
Publication venue: Bentham Open
Publication date
Field of study

In this article we have assembled the experience obtained from our participation in the imageCLEF evaluation task over the past two years. Exploitation on the use of linear combinations for image retrieval has been attempted by combining visual and textual sources of images. From our experiments we conclude that a mixed retrieval technique that applies both textual and visual retrieval in an interchangeably repeated manner improves the performance while overcoming the scalability limitations of visual retrieval. In particular, the mean average precision (MAP) has increased from 0.01 to 0.15 and 0.087 for 2009 and 2010 data, respectively, when content-based image retrieval (CBIR) is performed on the top 1000 results from textual retrieval based on natural language processing (NLP)

Crossref

PubMed Central

Recommended from our members

Lessons learned from the CHiC and SBS interactive tracks: a wishlist for interactive IR evaluation

Author: Bogers Toine
Gäde Maria
Hall Mark
Petras Vivien
Skov Mette
Publication venue
Publication date: 01/01/2017
Field of study

Over the course of the past two decades, the Interactive Tracks at TREC and INEX have contributed greatly to our knowledge of how to run an interactive IR evaluation campaign. In this position paper, we add to this body of knowledge by taking stock of our own experiences and challenges in organizing the CHIC and SBS Interactive Tracks from 2013 to 2016 in the form of a list of important properties of any future IIR evaluation campaigns

Open Research Online (The Open University)

VBN

Spoken content retrieval: A survey of techniques and technologies

Author: Ani Nenkova
C A. Nenkova
K. Mckeown
Kathleen Mckeown
Publication venue: 'Now Publishers'
Publication date: 01/01/2012
Field of study

Speech media, that is, digital audio and video containing spoken content, has blossomed in recent years. Large collections are accruing on the Internet as well as in private and enterprise settings. This growth has motivated extensive research on techniques and technologies that facilitate reliable indexing and retrieval. Spoken content retrieval (SCR) requires the combination of audio and speech processing technologies with methods from information retrieval (IR). SCR research initially investigated planned speech structured in document-like units, but has subsequently shifted focus to more informal spoken content produced spontaneously, outside of the studio and in conversational settings. This survey provides an overview of the field of SCR encompassing component technologies, the relationship of SCR to text IR and automatic speech recognition and user interaction issues. It is aimed at researchers with backgrounds in speech technology or IR who are seeking deeper insight on how these fields are integrated to support research and development, thus addressing the core challenges of SCR

CiteSeerX

Crossref

Irish Universities

DCU Online Research Access Service

Overview of the wikipediaMM task at ImageCLEF 2008

Author: Kludas J.
Tsikrika T. (Theodora)
Publication venue: 'Springer Fachmedien Wiesbaden GmbH'
Publication date: 01/09/2009
Field of study

The wikipediaMM task provides a testbed for the system-oriented evaluation of ad-hoc retrieval from a large collection of Wikipedia images. It became a part of the ImageCLEF evaluation campaign in 2008 with the aim of investigating the use of visual and textual sources in combination for improving the retrieval performance. This paper presents an overview of the task¿s resources, topics, assessments, participants' approaches, and main results

CWI's Institutional Repository

A Corpus for Hybrid Question Answering Systems

Author: Grau Brigitte
Ligozat Anne-Laure
Publication venue: 'Association for Computing Machinery (ACM)'
Publication date: 01/01/2018
Field of study

International audienceQuestion answering has been the focus of a lot of researches and evaluation campaigns, either for text-based systems (TREC and CLEF evaluation campaigns for example), or for knowledge-based systems (QALD, BioASQ). Few systems have effectively combined both types of resources and methods in order to exploit the fruitful- ness of merging the two kinds of information repositories. The only evaluation QA track that focuses on hybrid QA is QALD since 2014. As it is a recent task, few annotated data are available (around 150 questions). In this paper, we present a question answering dataset that was constructed to develop and evaluate hybrid question an- swering systems. In order to create this corpus, we collected several textual corpora and augmented them with entities and relations of a knowledge base by retrieving paths in the knowledge base which allow to answer the questions. The resulting corpus contains 4300 question-answer pairs and 1600 have a true link with DBpedia

Crossref