☆ 4.6 Article

Cross-company defect prediction via semi-supervised clustering-based data filtering and MSTrA-based transfer learning

SOFT COMPUTING (2018)

期刊

SOFT COMPUTING

卷 22, 期 10, 页码 3461-3472

出版社

SPRINGER

DOI: 10.1007/s00500-018-3093-1

关键词

Cross-company defect prediction; Transfer learning; SSDBSCAN; Multi-source TrAdaBoost

类别

Computer Science, Artificial Intelligence Computer Science, Interdisciplinary Applications

资金

National Natural Science Foundation of China [61070013, 61300042, U1135005, 71401128]
Fundamental Research Funds for the Central Universities [2042014kf0272, 2014211020201]
Natural Science Foundation of HuBei [2011CDB072]

向作者/读者索取更多资源

Protocol

社区支持

Reagent

社区支持

摘要

Cross-company defect prediction (CCDP) is a practical way that trains a prediction model by exploiting one or multiple projects of a source company and then applies the model to a target company. Unfortunately, larger irrelevant cross-company (CC) data usually make it difficult to build a prediction model with high performance. On the other hand, brute force leveraging of CC data poorly related to within-company data may decrease the prediction model performance. To address such issues, we aim to provide an effective solution for CCDP. First, we propose a novel semi-supervised clustering-based data filtering method (i.e., SSDBSCAN filter) to filter out irrelevant CC data. Second, based on the filtered CC data, we for the first time introduce multi-source TrAdaBoost algorithm, an effective transfer learning method, into CCDP to import knowledge not from one but from multiple sources to avoid negative transfer. Experiments on 15 public datasets indicate that: (1) our proposed SSDBSCAN filter achieves better overall performance than compared data filtering methods; (2) our proposed CCDP approach achieves the best overall performance among all tested CCDP approaches; and (3) our proposed CCDP approach performs significantly better than with-company defect prediction models.

Cross-company defect prediction via semi-supervised clustering-based data filtering and MSTrA-based transfer learning

期刊

SOFT COMPUTING

出版社

SPRINGER

关键词

类别

资金

向作者/读者索取更多资源

Protocol

Reagent

作者

我是这篇论文的作者

评论

主要评分

次要评分

新颖性

重要性

科学严谨性

评价这篇论文

推荐

Cross-company defect prediction via semi-supervised clustering-based data filtering and MSTrA-based transfer learning

期刊

SOFT COMPUTING

出版社

SPRINGER

关键词

类别

资金

向作者/读者索取更多资源

Protocol

Reagent

作者

我是这篇论文的作者

评论

主要评分

次要评分

新颖性

重要性

科学严谨性

评价这篇论文

推荐

导出引文

分享论文