4.6 Article

Text categorization: past and present

期刊

ARTIFICIAL INTELLIGENCE REVIEW
卷 54, 期 4, 页码 3007-3054

出版社

SPRINGER
DOI: 10.1007/s10462-020-09919-1

关键词

Text categorization; Conventional methods; Fuzzy logic; Deep learning; Nature-inspired algorithms; Graph-based methods

资金

  1. DST

向作者/读者索取更多资源

Automatic text categorization involves sorting text documents into predefined categories using machine learning algorithms. Key factors in text categorization include word sense, semantic relationships, and different methods within the conventional approaches.
Automatic text categorization is the operation of sorting out the text documents into pre-defined text categories using some machine learning algorithms. Normally, it defines the most important approaches to organizing and making the use of a large volume of information exists in unstructured form. Nowadays, text categorization is becoming an extensively researched field of text mining and processing of languages. Word sense, semantic relationships among terms, text documents and categories are quite essential in order of enhancing the performances of categorization. Various surveys on text categorization have already been available which involve techniques of various text representation schemes to such extent but do not include several approaches that have been explored in text categorization over the standard techniques. Here, an exhaustive analysis of different text categorization approaches over the conventional approaches has been undertaken. This survey paper explores a wide variety of algorithms used for categorizing text documents and tries to assemble the existing works into three basic fields: conventional methods, fuzzy logic-based methods, deep learning-based methods. Further, conventional methods have been categorized into three fields: text categorization using handcrafted features, text categorization using nature-inspired algorithms and text categorization using graph-based methods. Furthermore, this survey provides a clear idea about the available libraries used for different algorithms, availability of datasets, categorization technologies explored in various non-Indian and Indian languages as well.

作者

我是这篇论文的作者
点击您的名字以认领此论文并将其添加到您的个人资料中。

评论

主要评分

4.6
评分不足

次要评分

新颖性
-
重要性
-
科学严谨性
-
评价这篇论文

推荐

暂无数据
暂无数据