☆ 4.7 Article

Modification of the random forest algorithm to avoid statistical dependence problems when classifying remote sensing imagery

COMPUTERS & GEOSCIENCES (2017)

Journal

COMPUTERS & GEOSCIENCES

Volume 103, Issue -, Pages 1-11

Publisher

PERGAMON-ELSEVIER SCIENCE LTD

DOI: 10.1016/j.cageo.2017.02.012

Keywords

Classification; Random forest; Object-based image analysis; Bagging; Statistical independence

Funding

Prometeo Project, Secretariat of Higher Education, Science, Technology and Innovation, Gobierno de Ecuador

Ask authors/readers for more resources

Protocol

Community support

Reagent

Community support

Abstract

Random forest is a classification technique widely used in remote sensing. One of its advantages is that it produces an estimation of classification accuracy based on the so called out-of-bag cross-validation method. It is usually assumed that such estimation is not biased and may be used instead of validation based on an external data-set or a cross-validation external to the algorithm. In this paper we show that this is not necessarily the case when classifying remote sensing imagery using training areas with several pixels or objects. According to our results, out-of-bag cross-validation clearly overestimates accuracy, both overall and per class. The reason is that, in a training patch, pixels or objects are not independent (from a statistical point of view) of each other; however, they are split by bootstrapping into in bag and out-of-bag as if they were really independent. We believe that putting whole patch, rather than pixels/objects, in one or the other set would produce a less biased out-of-bag cross-validation. To deal with the problem, we propose a modification of the random forest algorithm to split training patches instead of the pixels (or objects) that compose them. This modified algorithm does not overestimate accuracy and has no lower predictive capability than the original. When its results are validated with an external data-set, the accuracy is not different from that obtained with the original algorithm. We analysed three remote sensing images with different classification approaches (pixel and object based); in the three cases reported, the modification we propose produces a less biased accuracy estimation.

Modification of the random forest algorithm to avoid statistical dependence problems when classifying remote sensing imagery

Journal

COMPUTERS & GEOSCIENCES

Publisher

PERGAMON-ELSEVIER SCIENCE LTD

Keywords

Categories

Funding

Ask authors/readers for more resources

Protocol

Reagent

Authors

I am an author on this paper

Reviews

Primary Rating

Secondary Ratings

Novelty

Significance

Scientific rigor

Rate this paper

Recommended

Modification of the random forest algorithm to avoid statistical dependence problems when classifying remote sensing imagery

Journal

COMPUTERS & GEOSCIENCES

Publisher

PERGAMON-ELSEVIER SCIENCE LTD

Keywords

Categories

Funding

Ask authors/readers for more resources

Protocol

Reagent

Authors

I am an author on this paper

Reviews

Primary Rating

Secondary Ratings

Novelty

Significance

Scientific rigor

Rate this paper

Recommended

Export Citation

Share Paper