Location of Repository

Cooperative strategy for web data mining and cleaning

By Yuefeng Li, Chengqi Zhang and Shichao Zhang

Abstract

While the Internet and World Wide Web have put a huge volume of low-quality information at the easy access of an information gathering system, filtering out irrelevant information has become a big challenge. In this paper, a Web data mining and cleaning strategy for information gathering is proposed. A data-mining model is presented for the data that come from multiple agents. Using the model, a data-cleaning algorithm is then presented to eliminate irrelevant data. To evaluate the data-cleaning strategy, an interpretation is given for the mining model according to evidence theory. An experiment is also conducted to evaluate the strategy using Web data. The experimental results have shown that the proposed strategy is efficient and promising

Topics: 080704 Information Retrieval and Web Search, 080100 ARTIFICIAL INTELLIGENCE AND IMAGE PROCESSING, Web mining, information fusion, information agents
Publisher: Taylor & Francis
Year: 2003
DOI identifier: 10.1080/713827173
OAI identifier: oai:eprints.qut.edu.au:7050

Suggested articles

Preview


To submit an update or takedown request for this paper, please submit an Update/Correction/Removal Request.