☆ 4.7 Article

Text clustering with feature selection by using statistical data

IEEE TRANSACTIONS ON KNOWLEDGE AND DATA ENGINEERING (2008)

Journal

IEEE TRANSACTIONS ON KNOWLEDGE AND DATA ENGINEERING

Volume 20, Issue 5, Pages 641-652

Publisher

IEEE COMPUTER SOC

DOI: 10.1109/TKDE.2007.190740

Keywords

text clustering; text mining; chi(2) Statistic; feature selection; performance analysis

Ask authors/readers for more resources

Protocol

Community support

Reagent

Community support

Abstract

Feature, selection is an important method for improving the efficiency and accuracy of text categorization algorithms by removing redundant and irrelevant terms from the corpus. In this paper, we propose a new supervised feature selection method, named CHIR, which is based on the chi(2) statistic and new statistical data that can measure the positive term-category dependency. We also propose a new text clustering algorithm, named Text Clustering with Feature Selection (TCFS). TCFS can incorporate CHIR to identify relevant features (i.e., terms) iteratively, and the clustering becomes 6 learning process. We compared TCFS and the K-means clustering algorithm in combination with different feature selection methods for various real data sets. Our experimental results show that TCFS with CHIR has better clustering accuracy in terms of the F-measure and the purity.

Text clustering with feature selection by using statistical data

Journal

IEEE TRANSACTIONS ON KNOWLEDGE AND DATA ENGINEERING

Publisher

IEEE COMPUTER SOC

Keywords

Categories

Ask authors/readers for more resources

Protocol

Reagent

Authors

I am an author on this paper

Reviews

Primary Rating

Secondary Ratings

Novelty

Significance

Scientific rigor

Rate this paper

Recommended

Text clustering with feature selection by using statistical data

Journal

IEEE TRANSACTIONS ON KNOWLEDGE AND DATA ENGINEERING

Publisher

IEEE COMPUTER SOC

Keywords

Categories

Ask authors/readers for more resources

Protocol

Reagent

Authors

I am an author on this paper

Reviews

Primary Rating

Secondary Ratings

Novelty

Significance

Scientific rigor

Rate this paper

Recommended

Export Citation

Share Paper