3.8 Proceedings Paper

CluWords: Exploiting Semantic Word Clustering Representation for Enhanced Topic Modeling

出版社

ASSOC COMPUTING MACHINERY
DOI: 10.1145/3289600.3291032

关键词

Data Representation; Topic Modeling; Word Embedding

资金

  1. CAPES
  2. CNPq
  3. Finep
  4. Fapemig
  5. Mundiale
  6. Astrein
  7. project InWeb
  8. project MASWeb

向作者/读者索取更多资源

In this paper, we advance the state-of-the-art in topic modeling by means of a new document representation based on pre-trained word embeddings for non-probabilistic matrix factorization. Specifically, our strategy, called CluWords, exploits the nearest words of a given pre-trained word embedding to generate meta-words capable of enhancing the document representation, in terms of both, syntactic and semantic information. The novel contributions of our solution include: (i) the introduction of a novel data representation for topic modeling based on syntactic and semantic relationships derived from distances calculated within a pre-trained word embedding space and (ii) the proposal of a new TF-IDF-based strategy, particularly developed to weight the CluWords. In our extensive experimentation evaluation, covering 12 datasets and 8 state-ofthe-art baselines, we exceed (with a few ties) in almost cases, with gains of more than 50% against the best baselines (achieving up to 80% against some runner-ups). Finally, we show that our method is able to improve document representation for the task of automatic text classification.

作者

我是这篇论文的作者
点击您的名字以认领此论文并将其添加到您的个人资料中。

评论

主要评分

3.8
评分不足

次要评分

新颖性
-
重要性
-
科学严谨性
-
评价这篇论文

推荐

暂无数据
暂无数据