Search CORE

2,304 research outputs found

BigFCM: Fast, Precise and Scalable FCM on Hadoop

Author: Ghadiri Nasser
Ghaffari Meysam
Nikbakht Mohammad Amin
Publication venue: 'Elsevier BV'
Publication date: 10/05/2016
Field of study

Clustering plays an important role in mining big data both as a modeling technique and a preprocessing step in many data mining process implementations. Fuzzy clustering provides more flexibility than non-fuzzy methods by allowing each data record to belong to more than one cluster to some degree. However, a serious challenge in fuzzy clustering is the lack of scalability. Massive datasets in emerging fields such as geosciences, biology and networking do require parallel and distributed computations with high performance to solve real-world problems. Although some clustering methods are already improved to execute on big data platforms, but their execution time is highly increased for large datasets. In this paper, a scalable Fuzzy C-Means (FCM) clustering named BigFCM is proposed and designed for the Hadoop distributed data platform. Based on the map-reduce programming model, it exploits several mechanisms including an efficient caching design to achieve several orders of magnitude reduction in execution time. Extensive evaluation over multi-gigabyte datasets shows that BigFCM is scalable while it preserves the quality of clustering

arXiv.org e-Print Archive

Western Sydney ResearchDirect

Applying subclustering and Lp distance in Weighted K-Means with distributed centroids

Author: Aggarwal
Baldi
Ball
Chan
Chatzis
Chen
Chiang
Cordeiro de Amorim
Hsu
Huang
Hubert
Jain
Jain
Ji
Ji
Kim
MacCuish
Makarenkov
Renato Cordeiro de Amorim
Rousseeuw
Steinley
Vladimir Makarenkov
Publication venue: 'Elsevier BV'
Publication date: 17/08/2015
Field of study

We consider the Weighted K-Means algorithm with distributed centroids aimed at clustering data sets with numerical, categorical and mixed types of data. Our approach allows given features (i.e., variables) to have different weights at different clusters. Thus, it supports the intuitive idea that features may have different degrees of relevance at different clusters. We use the Minkowski metric in a way that feature weights become feature re-scaling factors for any considered exponent. Moreover, the traditional Silhouette clustering validity index was adapted to deal with both numerical and categorical types of features. Finally, we show that our new method usually outperforms traditional K-Means as well as the recently proposed WK-DC clustering algorithm.Peer reviewe

University of Essex Research Repository

Crossref

University of Hertfordshire Research Archive

A hybrid supervised/unsupervised machine learning approach to solar flare prediction

Author: Benvenuto Federico
Campi Cristina
Massone Anna Maria
Piana Michele
Publication venue: 'American Astronomical Society'
Publication date: 21/06/2017
Field of study

We introduce a hybrid approach to solar flare prediction, whereby a supervised regularization method is used to realize feature importance and an unsupervised clustering method is used to realize the binary flare/no-flare decision. The approach is validated against NOAA SWPC data

arXiv.org e-Print Archive

Archivio istituzionale della ricerca - Università di Genova

Archivio istituzionale della ricerca - Università di Padova

MaxMin Linear Initialization for Fuzzy C-Means

Author: AM Bensaid
D Steinley
DJ Hand
EH Ruspini
GN Lance
HS Park
J. C. Dunn
JC Bezdek
ME Celebi
MJ Norušis
NR Pal
S Wold
T Caliński
T Su
TF Gonzalez
V Faber
W Wang
XL Xie
Publication venue: 'Springer Fachmedien Wiesbaden GmbH'
Publication date: 14/07/2018
Field of study

International audienceClustering is an extensive research area in data science. The aim of clustering is to discover groups and to identify interesting patterns in datasets. Crisp (hard) clustering considers that each data point belongs to one and only one cluster. However, it is inadequate as some data points may belong to several clusters, as is the case in text categorization. Thus, we need more flexible clustering. Fuzzy clustering methods, where each data point can belong to several clusters, are an interesting alternative. Yet, seeding iterative fuzzy algorithms to achieve high quality clustering is an issue. In this paper, we propose a new linear and efficient initialization algorithm MaxMin Linear to deal with this problem. Then, we validate our theoretical results through extensive experiments on a variety of numerical real-world and artificial datasets. We also test several validity indices, including a new validity index that we propose, Transformed Standardized Fuzzy Difference (TSFD)

arXiv.org e-Print Archive

Crossref

HAL

Hal-Diderot