☆ 4.6 Article

Theoretical measures of relative performance of classifiers for high dimensional data with small sample sizes

JOURNAL OF THE ROYAL STATISTICAL SOCIETY SERIES B-STATISTICAL METHODOLOGY (2008)

Journal

JOURNAL OF THE ROYAL STATISTICAL SOCIETY SERIES B-STATISTICAL METHODOLOGY

Volume 70, Issue -, Pages 159-173

Publisher

WILEY-BLACKWELL

DOI: 10.1111/j.1467-9868.2007.00631.x

Keywords

classification boundary; detection; distance-based classification; distance-weighted discrimination; higher criticism; nearest neighbour method; sparsity; support vector machine; thresholding; truncation

Ask authors/readers for more resources

Protocol

Community support

Reagent

Community support

Abstract

We suggest a technique, related to the concept of 'detection boundary' that was developed by Ingster and by Donoho and Jin, for comparing the theoretical performance of classifiers constructed from small training samples of very large vectors. The resulting 'classification boundaries' are obtained for a variety of distance-based methods, including the support vector machine, distance-weighted discrimination and kth-nearest-neighbour classifiers, for thresholded forms of those methods, and for techniques based on Donoho and Jin's higher criticism approach to signal detection. Assessed in these terms, standard distance-based methods are shown to be capable only of detecting differences between populations when those differences can be estimated consistently. However, the thresholded forms of distance-based classifiers can do better, and in particular can correctly classify data even when differences between distributions are only detectable, not estimable. Other methods, including higher criticism classifiers, can on occasion perform better still, but they tend to be more limited in scope, requiring substantially more information about the marginal distributions. Moreover, as tail weight becomes heavier the classification boundaries of methods designed for particular distribution types can converge to, and achieve, the boundary for thresholded nearest neighbour approaches. For example, although higher criticism has a lower classification boundary, and in this sense performs better, in the case of normal data, the boundaries are identical for exponentially distributed data when both sample sizes equal 1.

Theoretical measures of relative performance of classifiers for high dimensional data with small sample sizes

Journal

JOURNAL OF THE ROYAL STATISTICAL SOCIETY SERIES B-STATISTICAL METHODOLOGY

Publisher

WILEY-BLACKWELL

Keywords

Categories

Ask authors/readers for more resources

Protocol

Reagent

Authors

I am an author on this paper

Reviews

Primary Rating

Secondary Ratings

Novelty

Significance

Scientific rigor

Rate this paper

Recommended

Theoretical measures of relative performance of classifiers for high dimensional data with small sample sizes

Journal

JOURNAL OF THE ROYAL STATISTICAL SOCIETY SERIES B-STATISTICAL METHODOLOGY

Publisher

WILEY-BLACKWELL

Keywords

Categories

Ask authors/readers for more resources

Protocol

Reagent

Authors

I am an author on this paper

Reviews

Primary Rating

Secondary Ratings

Novelty

Significance

Scientific rigor

Rate this paper

Recommended

Export Citation

Share Paper