☆ 4.8 Article

A novel improved random forest for text classification using feature ranking and optimal number of trees

JOURNAL OF KING SAUD UNIVERSITY-COMPUTER AND INFORMATION SCIENCES (2022)

期刊

JOURNAL OF KING SAUD UNIVERSITY-COMPUTER AND INFORMATION SCIENCES

卷 34, 期 6, 页码 2733-2742

出版社

ELSEVIER

DOI: 10.1016/j.jksuci.2022.03.012

关键词

Improved random forest; Text classification; Feature ranking; Decision tree optimization; Machine learning; Feature reduction

类别

Computer Science, Information Systems

向作者/读者索取更多资源

Protocol

社区支持

Reagent

社区支持

智能总结 New
摘要

This study presents an improved random forest model for text classification, called IRFTC, which incorporates bootstrapping and random subspace methods. It achieves better classification performance and outperforms other machine learning models.

Machine learning-based models like random forest (RF) have been widely deployed in diverse domains such as image processing, health care, and text processing, etc. during the past few years. The RF is a prominent technique for handling imbalanced data and performs significantly better than other machine learning models due to its parallel architecture. This study presents an improved random forest for text classification, called improved random forest for text classification (IRFTC), that incorporates bootstrapping and random subspace methods simultaneously. The IRFTC removes unimportant (less important) features, adds a number of trees in the forest on each iteration, and monitors the classification performance of RF. Classification accuracy is determined with respect to the number of trees which defines the optimal number of trees for IRFTC. Feature ranking is determined using the quality of the split in a tree. The proposed IRFTC is applied on four different benchmark datasets, binary and multiclass, to validate its performance in this study. Results indicate that IRFTC outperforms both the traditional RF, as well as, other machine learning models such as logistic regression, support vector machine, Naive Bayes, and decision trees.

A novel improved random forest for text classification using feature ranking and optimal number of trees

期刊

JOURNAL OF KING SAUD UNIVERSITY-COMPUTER AND INFORMATION SCIENCES

出版社

ELSEVIER

关键词

类别

向作者/读者索取更多资源

Protocol

Reagent

作者

我是这篇论文的作者

评论

主要评分

次要评分

新颖性

重要性

科学严谨性

评价这篇论文

推荐

A novel improved random forest for text classification using feature ranking and optimal number of trees

期刊

JOURNAL OF KING SAUD UNIVERSITY-COMPUTER AND INFORMATION SCIENCES

出版社

ELSEVIER

关键词

类别

向作者/读者索取更多资源

Protocol

Reagent

作者

我是这篇论文的作者

评论

主要评分

次要评分

新颖性

重要性

科学严谨性

评价这篇论文

推荐

导出引文

分享论文