☆ 4.5 Article

DOLPHIN: An Efficient Algorithm for Mining Distance-Based Outliers in Very Large Datasets

ACM TRANSACTIONS ON KNOWLEDGE DISCOVERY FROM DATA (2009)

Journal

ACM TRANSACTIONS ON KNOWLEDGE DISCOVERY FROM DATA

Volume 3, Issue 1, Pages -

Publisher

ASSOC COMPUTING MACHINERY

DOI: 10.1145/1497577.1497581

Keywords

Data mining; outlier detection; distance-based outliers

Ask authors/readers for more resources

Protocol

Community support

Reagent

Community support

Abstract

In this work a novel distance-based outlier detection algorithm, named DOLPHIN, working on disk-resident datasets and whose I/O cost corresponds to the cost of sequentially reading the input dataset file twice, is presented. It is both theoretically and empirically shown that the main memory usage of DOLPHIN amounts to a small fraction of the dataset and that DOLPHIN has linear time performance with respect to the dataset size. DOLPHIN gains efficiency by naturally merging together in a unified schema three strategies, namely the selection policy of objects to be maintained in main memory, usage of pruning rules, and similarity search techniques. Importantly, similarity search is accomplished by the algorithm without the need of preliminarily indexing the whole dataset, as other methods do. The algorithm is simple to implement and it can be used with any type of data, belonging to either metric or nonmetric spaces. Moreover, a modification to the basic method allows DOLPHIN to deal with the scenario in which the available buffer of main memory is smaller than its standard requirements. DOLPHIN has been compared with state-of-the-art distance-based outlier detection algorithms, showing that it is much more efficient.

DOLPHIN: An Efficient Algorithm for Mining Distance-Based Outliers in Very Large Datasets

Journal

ACM TRANSACTIONS ON KNOWLEDGE DISCOVERY FROM DATA

Publisher

ASSOC COMPUTING MACHINERY

Keywords

Categories

Ask authors/readers for more resources

Protocol

Reagent

Authors

I am an author on this paper

Reviews

Primary Rating

Secondary Ratings

Novelty

Significance

Scientific rigor

Rate this paper

Recommended

DOLPHIN: An Efficient Algorithm for Mining Distance-Based Outliers in Very Large Datasets

Journal

ACM TRANSACTIONS ON KNOWLEDGE DISCOVERY FROM DATA

Publisher

ASSOC COMPUTING MACHINERY

Keywords

Categories

Ask authors/readers for more resources

Protocol

Reagent

Authors

I am an author on this paper

Reviews

Primary Rating

Secondary Ratings

Novelty

Significance

Scientific rigor

Rate this paper

Recommended

Export Citation

Share Paper