☆ 4.6 Article

Neural Network-based Detection of Self-Admitted Technical Debt: From Performance to Explainability

ACM TRANSACTIONS ON SOFTWARE ENGINEERING AND METHODOLOGY (2019)

期刊

ACM TRANSACTIONS ON SOFTWARE ENGINEERING AND METHODOLOGY

卷 28, 期 3, 页码 -

出版社

ASSOC COMPUTING MACHINERY

DOI: 10.1145/3324916

关键词

Self-admitted technical debt; convolutional neural network; cross project prediction; model explainability; model generalizability; model adaptability

类别

Computer Science, Software Engineering

资金

National Key Research and Development Program of China [2018YFB1003904]
NSFC Program [61602403]
Project of Science and Technology Research and Development Program of China Railway Corporation [P2018X002]
Fundamental Research Funds for the Central Universities

向作者/读者索取更多资源

Protocol

社区支持

Reagent

社区支持

摘要

Technical debt is a metaphor to reflect the tradeoff software engineers make between short-term benefits and long-term stability. Self-admitted technical debt (SKID), a variant of technical debt, has been proposed to identify debt that is Intentionally introduced during software development, e.g., temporary fixes and workarounds. Previous studies have leveraged human-summarized patterns (which represent n-gram phrases that can be used to identify SATD) or text-mining techniques to detect SATD in source code comments. However, several characteristics of SAID features in code comments, such as vocabulary diversity, project uniqueness, length, and semantic variations, pose a big challenge to the accuracy of pattern or traditional text-mining-based SAID detection, especially for cross-project deployment. Furthermore, although traditional text-mining-based method outperforms pattern-based method in prediction accuracy, the text features it uses are less intuitive than human-summarized patterns, which makes the prediction results hard to explain. To improve the accuracy of SATD prediction, especially for cross-project prediction, we propose a Convolutional Neural Network- (CNN) based approach for classifying code comments as SATD or non-SATD. To improve the explainability of our model's prediction results, we exploit the computational structure of CNNs to identify key phrases and patterns in code comments that are most relevant to SATD. We have conducted an extensive set of experiments with 62,566 code comments from 10 open-source projects and a user study with 150 comments of another three projects. Our evaluation confirms the effectiveness of different aspects of our approach and its superior performance, generalizability, adaptability, and explainability over current state-of-the-art traditional text-mining-based methods for SATD classification.

Neural Network-based Detection of Self-Admitted Technical Debt: From Performance to Explainability

期刊

ACM TRANSACTIONS ON SOFTWARE ENGINEERING AND METHODOLOGY

出版社

ASSOC COMPUTING MACHINERY

关键词

类别

资金

向作者/读者索取更多资源

Protocol

Reagent

作者

我是这篇论文的作者

评论

主要评分

次要评分

新颖性

重要性

科学严谨性

评价这篇论文

推荐

Neural Network-based Detection of Self-Admitted Technical Debt: From Performance to Explainability

期刊

ACM TRANSACTIONS ON SOFTWARE ENGINEERING AND METHODOLOGY

出版社

ASSOC COMPUTING MACHINERY

关键词

类别

资金

向作者/读者索取更多资源

Protocol

Reagent

作者

我是这篇论文的作者

评论

主要评分

次要评分

新颖性

重要性

科学严谨性

评价这篇论文

推荐

导出引文

分享论文