☆ 4.5 Article

A pseudo-nearest-neighbor approach for missing data recovery on Gaussian random data sets

PATTERN RECOGNITION LETTERS (2002)

Journal

PATTERN RECOGNITION LETTERS

Volume 23, Issue 13, Pages 1613-1622

Publisher

ELSEVIER SCIENCE BV

DOI: 10.1016/S0167-8655(02)00125-3

Keywords

missing data; missing data recovery; data imputation; data clustering; Gaussian data distribution; data mining

Ask authors/readers for more resources

Protocol

Community support

Reagent

Community support

Abstract

Missing data handling is an important preparation step for most data discrimination or mining tasks. Inappropriate treatment of missing data may cause large errors or false results. In this paper, we study the effect of a missing data recovery method, namely the pseudo-nearest-neighbor substitution approach, on Gaussian distributed data sets that represent typical cases in data discrimination and data mining applications. The error rate of the proposed recovery method is evaluated by comparing the clustering results of the recovered data sets to the clustering results obtained on the originally complete data sets. The results are also compared with that obtained by applying two other missing data handling methods, the constant default value substitution and the missing data ignorance (non-substitution) methods. The experiment results provided a valuable insight to the improvement of the accuracy for data discrimination and knowledge discovery on large data sets containing missing values. (C) 2002 Elsevier Science B.V. All rights reserved.

A pseudo-nearest-neighbor approach for missing data recovery on Gaussian random data sets

Journal

PATTERN RECOGNITION LETTERS

Publisher

ELSEVIER SCIENCE BV

Keywords

Categories

Ask authors/readers for more resources

Protocol

Reagent

Authors

I am an author on this paper

Reviews

Primary Rating

Secondary Ratings

Novelty

Significance

Scientific rigor

Rate this paper

Recommended

A pseudo-nearest-neighbor approach for missing data recovery on Gaussian random data sets

Journal

PATTERN RECOGNITION LETTERS

Publisher

ELSEVIER SCIENCE BV

Keywords

Categories

Ask authors/readers for more resources

Protocol

Reagent

Authors

I am an author on this paper

Reviews

Primary Rating

Secondary Ratings

Novelty

Significance

Scientific rigor

Rate this paper

Recommended

Export Citation

Share Paper