A NEW ANOMALOUS TEXT DETECTION APPROACH USING UNSUPERVISED METHODS

Amouee, Elham; Bahaghighat, Mahdi; Ghorbani, Mohsen; Zanjireh, Morteza Mohammadi

A NEW ANOMALOUS TEXT DETECTION APPROACH USING UNSUPERVISED METHODS

Authors: Elham Amouee
Mahdi Bahaghighat
Mohsen Ghorbani
Morteza Mohammadi Zanjireh
Publication date: 8 October 2020
Publisher: Published by the University of Niš, Serbia

Abstract

Increasing size of text data in databases requires appropriate classiﬁcation and analysis in order to acquire knowledge and improve the quality of decision-making in organizations. The process of discovering the hidden patterns in the data set, called data mining, requires access to quality data in order to receive a valid response from the system. Detecting and removing anomalous data is one of the pre-processing steps and cleaning data in this process. Methods for anomalous data detection are generally classiﬁed into three groups including supervised, semi-supervised, and unsupervised. This research tried to oﬀer an unsupervised approach for spotting the anomalous data in text collections. In the proposed method, a combination of two approaches (i.e., clustering-based and distance-based) is used for detecting anomaly in the text data. In order to evaluate the eﬃciency of the proposed approach, this method is applied on four labeled data sets. The accuracy of Na¨ıve Bayes classiﬁcation algorithms and decision tree are compared before and after removal of anomalous data with the proposed method and some other methods such as Density-based spatial clustering of applications with noise (DBSCAN). Our proposed method shows that accuracy of more than 92.39% can be achieved. In general, the results revealed that in most cases the proposed method has a good performance

Similar works

Full text

Open in the Core reader

Download PDF

Available Versions

University of Niš: Facta Universitatis (E-Journals) / Универзитет у Нишу

oai:casopisi.junis.ni.ac.rs:ar...

Last time updated on 30/11/2020