☆ 4.5 Article

A systematic and comprehensive investigation of methods to build and evaluate fault prediction models

JOURNAL OF SYSTEMS AND SOFTWARE (2010)

Journal

JOURNAL OF SYSTEMS AND SOFTWARE

Volume 83, Issue 1, Pages 2-17

Publisher

ELSEVIER SCIENCE INC

DOI: 10.1016/j.jss.2009.06.055

Keywords

Fault prediction models; Cost-effectiveness; Verification

Ask authors/readers for more resources

Protocol

Community support

Reagent

Community support

Abstract

This paper describes a study performed in an industrial setting that attempts to build predictive models to identify parts of a Java system with a high fault probability. The system under consideration is constantly evolving as several releases a year are shipped to customers. Developers usually have limited resources for their testing and would like to devote extra resources to faulty system parts. The main research focus of this paper is to systematically assess three aspects on how to build and evaluate fault-proneness models in the context of this large Java legacy system development project: (1) compare many data mining and machine learning techniques to build fault-proneness models, (2) assess the impact of using different metric sets such as source code structural measures and change/fault history (process measures), and (3) compare several alternative ways of assessing the performance of the models, in terms of (i) confusion matrix criteria such as accuracy and precision/recall, (ii) ranking ability, using the receiver operating characteristic area (ROC), and (iii) our proposed cost-effectiveness measure (CE). The results of the study indicate that the choice of fault-proneness modeling technique has limited impact on the resulting classification accuracy or cost-effectiveness. There is however large differences between the individual metric sets in terms of cost-effectiveness, and although the process measures are among the most expensive ones to collect, including them as candidate measures significantly improves the prediction models compared with models that only include structural measures and/or their deltas between releases - both in terms of ROC area and in terms of CE. Further, we observe that what is considered the best model is highly dependent on the criteria that are used to evaluate and compare the models. And the regular confusion matrix criteria, although popular, are not clearly related to the problem at hand, namely the cost-effectiveness of using fault-proneness prediction models to focus verification efforts to deliver software with less faults at less cost. (C) 2009 Elsevier Inc. All rights reserved.

A systematic and comprehensive investigation of methods to build and evaluate fault prediction models

Journal

JOURNAL OF SYSTEMS AND SOFTWARE

Publisher

ELSEVIER SCIENCE INC

Keywords

Categories

Ask authors/readers for more resources

Protocol

Reagent

Authors

I am an author on this paper

Reviews

Primary Rating

Secondary Ratings

Novelty

Significance

Scientific rigor

Rate this paper

Recommended

A systematic and comprehensive investigation of methods to build and evaluate fault prediction models

Journal

JOURNAL OF SYSTEMS AND SOFTWARE

Publisher

ELSEVIER SCIENCE INC

Keywords

Categories

Ask authors/readers for more resources

Protocol

Reagent

Authors

I am an author on this paper

Reviews

Primary Rating

Secondary Ratings

Novelty

Significance

Scientific rigor

Rate this paper

Recommended

Export Citation

Share Paper