4.5 Article

Smoothing target encoding and class center-based firefly algorithm for handling missing values in categorical variable

期刊

JOURNAL OF BIG DATA
卷 10, 期 1, 页码 -

出版社

SPRINGERNATURE
DOI: 10.1186/s40537-022-00679-z

关键词

Missing data; Encoding; Smoothing; Firefly algorithm; Class center

向作者/读者索取更多资源

This study aims to perform the imputation process using the smoothing target encoding (STE) method combined with C3FA and standard deviation (STD), and compare it with other imputation methods. The results showed that the proposed method (C3FA-STD) achieved AUC, CA, F1-Score, precision, and recall values of 0.939, 0.882, 0.881, 0.881, and 0.882, respectively, based on the evaluation using the kNN classifier.
One of the most common causes of incompleteness is missing data, which occurs when no data value for the variables in observation is stored. An adaptive approach model outperforming other numerical methods in the classification problem was developed using the class center-based Firefly algorithm by incorporating attribute correlations into the imputation process (C3FA). However, this model has not been tested on categorical data, which is essential in the preprocessing stage. Encoding is used to convert text or Boolean values in categorical data into numeric parameters, and the target encoding method is often utilized. This method uses target variable information to encode categorical data and it carries the risk of overfitting and inaccuracy within the infrequent categories. This study aims to use the smoothing target encoding (STE) method to perform the imputation process by combining C3FA and standard deviation (STD) and compare by several imputation methods. The results on the tic tac toe dataset showed that the proposed method (C3FA-STD) produced AUC, CA, F1-Score, precision, and recall values of 0.939, 0.882, 0.881, 0.881, and 0.882, respectively, based on the evaluation using the kNN classifier.

作者

我是这篇论文的作者
点击您的名字以认领此论文并将其添加到您的个人资料中。

评论

主要评分

4.5
评分不足

次要评分

新颖性
-
重要性
-
科学严谨性
-
评价这篇论文

推荐

暂无数据
暂无数据