期刊
BIG DATA MINING AND ANALYTICS
卷 4, 期 3, 页码 183-194出版社
TSINGHUA UNIV PRESS
DOI: 10.26599/BDMA.2021.9020001
关键词
density-based clustering; incomplete data; clustering algorihtm
资金
- National Natural Science Foundation of China [U1866602, 71773025]
- National Key Research and Development Program of China [2020YFB1006104]
Density-based clustering is an important category, but traditional techniques for handling missing values are not suitable. A novel approach based on Bayesian theory, which conducts imputation and clustering concurrently, shows effectiveness in experiments.
Density-based clustering is an important category among clustering algorithms. In real applications, many datasets suffer from incompleteness. Traditional imputation technologies or other techniques for handling missing values are not suitable for density-based clustering and decrease clustering result quality. To avoid these problems, we develop a novel density-based clustering approach for incomplete data based on Bayesian theory, which conducts imputation and clustering concurrently and makes use of intermediate clustering results. To avoid the impact of low-density areas inside non-convex clusters, we introduce a local imputation clustering algorithm, which aims to impute points to high-density local areas. The performances of the proposed algorithms are evaluated using ten synthetic datasets and five real-world datasets with induced missing values. The experimental results show the effectiveness of the proposed algorithms.
作者
我是这篇论文的作者
点击您的名字以认领此论文并将其添加到您的个人资料中。
推荐
暂无数据