Search CORE

4 research outputs found

Clustering Weblogs on the Basis of a Topic Detection Method

Author: Cardiff John
Perez-Tellez Fernando
Pinto David
Rosso Paolo
Publication venue: Dublin Institute of Technology
Publication date: 01/01/2010
Field of study

In recent years we have seen a vast increase in the volume of information published on weblog sites and also the creation of new web technologies where people discuss actual events. The need for automatic tools to organize this massive amount of information is clear, but the particular characteristics of weblogs such as shortness and overlapping vocabulary make this task difficult. In this work, we present a novel methodology to cluster weblog posts according to the topics discussed therein. This methodology is based on a generative probabilistic model in conjunction with a Self-Term Expansion methodology. We present our results which demonstrate a considerable improvement over the baseline

Crossref

Arrow@TUDublin

Extracting common emotions from blogs based on fine-grained sentiment clustering

Author: FENG Shi
GAO Wei
WANG Daling
WONG Kam-Fai
YU Ge
Publication venue: 'Springer Science and Business Media LLC'
Publication date: 01/07/2010
Field of study

Crossref

Institutional Knowledge at Singapore Management University

Prototype/topic based Clustering Method for Weblogs

Author: Cardiff John
Perez-Tellez Fernando
Pinto Avendaño David Eduardo
Rosso Paolo
Publication venue: 'IOS Press'
Publication date: 01/01/2016
Field of study

[EN] In the last 10 years, the information generated on weblog sites has increased exponentially, resulting in a clear need for intelligent approaches to analyse and organise this massive amount of information. In this work, we present a methodology to cluster weblog posts according to the topics discussed therein, which we derive by text analysis. We have called the methodology Prototype/Topic Based Clustering, an approach which is based on a generative probabilistic model in conjunction with a Self-Term Expansion methodology. The usage of the Self-Term Expansion methodology is to improve the representation of the data and the generative probabilistic model is employed to identify relevant topics discussed in the weblogs. We have modified the generative probabilistic model in order to exploit predefined initialisations of the model and have performed our experiments in narrow and wide domain subsets. The results of our approach have demonstrated a considerable improvement over the pre-defined baseline and alternative state of the art approaches, achieving an improvement of up to 20% in many cases. The experiments were performed on both narrow and wide domain datasets, with the latter showing better improvement. However in both cases, our results outperformed the baseline and state of the art algorithms.The work of the third author was carried out in the framework of the WIQ-EI IRSES project (Grant No. 269180) within the FP7 Marie Curie, the DIANA APPLICATIONS Finding Hidden Knowledge in Texts: Applications (TIN2012-38603-C02-01) project and the VLC/CAMPUS Microcluster on Multimodal Interaction in Intelligent Systems.Perez-Tellez, F.; Cardiff, J.; Rosso, P.; Pinto Avendaño, DE. (2016). Prototype/topic based Clustering Method for Weblogs. Intelligent Data Analysis. 20(1):47-65. https://doi.org/10.3233/IDA-150793S476520

Crossref

RiuNet

Clustering blogs with collective wisdom

Author: Huan Liu
Magdiel Galan
Nitin Agarwal
Shankar Subramanya
Publication venue
Publication date: 01/01/2008
Field of study

Blogosphere is expanding in an unprecedented speed. A better understanding of the blogosphere can greatly facilitate the development of the Social Web to serve the needs of users, service providers and advertisers. One important task in this process is clustering blog sites. Clustering blog sites presents new challenges. We propose to tap into collective wisdom in clustering blog sites, present statistical and visual results, report findings, and suggest future work extending to many real-world applications.

CiteSeerX