Search CORE

2,032 research outputs found

Setting per-field normalisation hyper-parameters for the named-page finding search task

Author: He B.
Ounis I.
Publication venue: 'Springer Science and Business Media LLC'
Publication date: 01/01/2007
Field of study

Per-field normalisation has been shown to be effective for Web search tasks, e.g. named-page finding. However, per-field normalisation also suffers from having hyper-parameters to tune on a per-field basis. In this paper, we argue that the purpose of per-field normalisation is to adjust the linear relationship between field length and term frequency. We experiment with standard Web test collections, using three document fields, namely the body of the document, its title, and the anchor text of its incoming links. From our experiments, we find that across different collections, the linear correlation values, given by the optimised hyper-parameter settings, are proportional to the maximum negative linear correlation. Based on this observation, we devise an automatic method for setting the per-field normalisation hyper-parameter values without the use of relevance assessment for tuning. According to the evaluation results, this method is shown to be effective for the body and title fields. In addition, the difficulty in setting the per-field normalisation hyper-parameter for the anchor text field is explained

CiteSeerX

Enlighten

Ensemble clustering for result diversification

Author: Hiemstra Djoerd
Nguyen Dong-Phuong
Publication venue: NIST
Publication date: 01/01/2012
Field of study

This paper describes the participation of the University of Twente in the Web track of TREC 2012. Our baseline approach uses the Mirex toolkit, an open source tool that sequantially scans all the documents. For result diversification, we experimented with improving the quality of clusters through ensemble clustering. We combined clusters obtained by different clustering methods (such as LDA and K-means) and clusters obtained by using different types of data (such as document text and anchor text). Our two-layer ensemble run performed better than the LDA based diversification and also better than a non-diversification run

Edinburgh Research Explorer

Radboud Repository

University of Twente Research Information

Rhetorical Classification of Anchor Text for Citation Recommendation

Author: Clare Amanda
Duma Daniel
Klein Ewan
Liakata Maria
Ravenscroft James
Publication venue
Publication date: 01/09/2016
Field of study

Crossref

Aberystwyth Research Portal

Edinburgh Research Explorer

Web Page Retrieval by Combining Evidence

Author: Alonso-Berrocal José-Luis
G.-Figuerola Carlos
Rodríguez-Vázquez-de-Aldana Emilio
Zazo Ángel F.
Publication venue
Publication date: 01/01/2006
Field of study

The participation of the REINA Research Group in WebCLEF 2005 focused in the monolingual mixed task. Queries or topics are of two types: named and home pages. For both, we first perform a search by thematic contents; for the same query, we do a search in several elements of information from every page (title, some meta tags, anchor text) and then we combine the results. For queries about home pages, we try to detect using a method based in some keywords and their patterns of use. After, a re-rank of the results of the thematic contents retrieval is performed, based on Page-Rank and Centrality coeficients

E-LIS

Venue Recommendation and Web Search Based on Anchor Text

Author: Hashemi S.H.
Kamps J.
Publication venue: National Institute for Standards and Technology
Publication date: 01/01/2014
Field of study

International Migration, Integration and Social Cohesion online publications

UvA-DARE

Temporal anchor text as proxy for real user queries

Author: Samar T. (Thaer)
Vries A.P. (Arjen) de
Publication venue
Publication date: 01/01/2015
Field of study

CWI's Institutional Repository

Coping with noise in a real-world weblog crawler and retrieval system

Author: Ferguson Paul
Lanagan James
O'Hare Neil
Smeaton Alan F.
Publication venue
Publication date: 01/05/2010
Field of study

In this paper we examine the effects of noise when creating a real-world weblog corpus for information retrieval. We focus on the DiffPost (Lee et al. 2008) approach to noise removal from blog pages, examining the difficulties encountered when crawling the blogosphere during the creation of a real-world corpus of blog pages. We introduce and evaluate a number of enhancements to the original DiffPost approach in order to increase the robustness of the algorithm. We then extend DiffPost by looking at the anchor-text to text ratio, and dis- cover that the time-interval between crawls is more impor- tant to the successful application of noise-removal algorithms within the blog context, than any additional improvements to the removal algorithm itself

Irish Universities

DCU Online Research Access Service

Lost but not forgotten: finding pages on the unarchived web

Author: Ben-David A.
Huurdeman H.C.
Kamps J.
Rogers R.A. (Richard)
Samar T. (Thaer)
Vries A.P. (Arjen) de
Publication venue: 'Springer Science and Business Media LLC'
Publication date: 01/01/2015
Field of study

Web archives attempt to preserve the fast changing web, yet they will always be incomplete. Due to restrictions in crawling depth, crawling frequency, and restrictive selection policies, large parts of the Web are unarchived and, therefore, lost to posterity. In this paper, we propose an approach to uncover unarchived web pages and websites and to reconstruct different types of descriptions for these pages and sites, based on links and anchor text in the set of crawled pages. We experiment with this approach on the Dutch Web Archive and evaluate the usefulness of page and host-level representations of unarchived content. Our main findings are the following: First, the crawled web contains evidence of a remarkable number of unarchived pages and websites, potentially dramatically increasing the coverage of a Web archive. Second, the link and anchor text have a highly skewed distribution: popular pages such as home pages have more links pointing to them and more terms in the anchor text, but the richness tapers off quickly. Aggregating web page evidence to the host-level leads to significantly richer representations, but the distribution remains skewed. Third, the succinct representation is generally rich enough to uniquely identify pages on the unarchived web: in a known-item search setting we can retrieve unarchived web pages within the first ranks on average, with host-level representations leading to further improvement of the retrieval effectiveness for websites

Crossref

CWI's Institutional Repository

Springer - Publisher Connector

UvA-DARE

International Migration, Integration and Social Cohesion online publications

Query reformulation using anchor text

Author
Publication venue: 'Association for Computing Machinery (ACM)'
Publication date: 01/01/2010
Field of study

Crossref