4.6 Article

A survey on data preprocessing for data stream mining: Current status and future directions

期刊

NEUROCOMPUTING
卷 239, 期 -, 页码 39-57

出版社

ELSEVIER
DOI: 10.1016/j.neucom.2017.01.078

关键词

Data mining; Data stream; Concept drift; Data preprocessing; Data reduction; Feature selection; Instance selection; Data discretization; Online learning

资金

  1. Spanish National Research [TIN2014-57251-P]
  2. Foundation BBVA [75/2016]
  3. Andalusian Research Plan [P11-TIC-7765]
  4. Polish National Science Center [DEC2013/09/B/ST6/02264]
  5. Spanish Ministry of Education and Science [FPU13/00047]

向作者/读者索取更多资源

Data preprocessing and reduction have become essential techniques in current knowledge discovery scenarios, dominated by increasingly large datasets. These methods aim at reducing the complexity inherent to real-world datasets, so that they can be easily processed by current data mining solutions. Advantages of such approaches include, among others, a faster and more precise learning process, and more understandable structure of raw data. However, in the context of data preprocessing techniques for data streams have a long road ahead of them, despite online learning is growing in importance thanks to the development of Internet and technologies for massive data collection. Throughout this survey, we summarize, categorize and analyze those contributions on data preprocessing that cope with streaming data. This work also takes into account the existing relationships between the different families of methods (feature and instance selection, and discretization). To enrich our study, we conduct thorough experiments using the most relevant contributions and present an analysis of their predictive performance, reduction rates, computational time, and memory usage. Finally, we offer general advices about existing data stream preprocessing algorithms, as well as discuss emerging future challenges to be faced in the domain of data stream preprocessing. (C) 2017 Elsevier B.V. All rights reserved.

作者

我是这篇论文的作者
点击您的名字以认领此论文并将其添加到您的个人资料中。

评论

主要评分

4.6
评分不足

次要评分

新颖性
-
重要性
-
科学严谨性
-
评价这篇论文

推荐

暂无数据
暂无数据