☆ 3.8 Article

A NEW ANOMALOUS TEXT DETECTION APPROACH USING UNSUPERVISED METHODS

FACTA UNIVERSITATIS-SERIES ELECTRONICS AND ENERGETICS (2020)

期刊

FACTA UNIVERSITATIS-SERIES ELECTRONICS AND ENERGETICS

卷 33, 期 4, 页码 631-653

出版社

UNIV NIS

DOI: 10.2298/FUEE2004631A

关键词

Anomaly detection; text mining; unsupervised learning; clustering; pre-processing; DBSCAN algorithm

类别

Engineering, Electrical & Electronic

向作者/读者索取更多资源

Protocol

社区支持

Reagent

社区支持

摘要

Increasing size of text data in databases requires appropriate classification and analysis in order to acquire knowledge and improve the quality of decision-making in organizations. The process of discovering the hidden patterns in the data set, called data mining, requires access to quality data in order to receive a valid response from the system. Detecting and removing anomalous data is one of the pre-processing steps and cleaning data in this process. Methods for anomalous data detection are generally classified into three groups including supervised, semi-supervised, and unsupervised. This research tried to offer an unsupervised approach for spotting the anomalous data in text collections. In the proposed method, a combination of two approaches (i.e., clustering-based and distance-based) is used for detecting anomaly in the text data. In order to evaluate the efficiency of the proposed approach, this method is applied on four labeled data sets. The accuracy of Naive Bayes classification algorithms and decision tree are compared before and after removal of anomalous data with the proposed method and some other methods such as Density-based spatial clustering of applications with noise (DBSCAN). Our proposed method shows that accuracy of more than 92.39% can he achieved. In general, the results revealed that in most cases the proposed method has a good performance.

A NEW ANOMALOUS TEXT DETECTION APPROACH USING UNSUPERVISED METHODS

期刊

FACTA UNIVERSITATIS-SERIES ELECTRONICS AND ENERGETICS

出版社

UNIV NIS

关键词

类别

向作者/读者索取更多资源

Protocol

Reagent

作者

我是这篇论文的作者

评论

主要评分

次要评分

新颖性

重要性

科学严谨性

评价这篇论文

推荐

A NEW ANOMALOUS TEXT DETECTION APPROACH USING UNSUPERVISED METHODS

期刊

FACTA UNIVERSITATIS-SERIES ELECTRONICS AND ENERGETICS

出版社

UNIV NIS

关键词

类别

向作者/读者索取更多资源

Protocol

Reagent

作者

我是这篇论文的作者

评论

主要评分

次要评分

新颖性

重要性

科学严谨性

评价这篇论文

推荐

导出引文

分享论文