☆ 4.7 Article

Fast Online EM for Big Topic Modeling

IEEE TRANSACTIONS ON KNOWLEDGE AND DATA ENGINEERING (2016)

期刊

IEEE TRANSACTIONS ON KNOWLEDGE AND DATA ENGINEERING

卷 28, 期 3, 页码 675-688

出版社

IEEE COMPUTER SOC

DOI: 10.1109/TKDE.2015.2492565

关键词

Latent Dirichlet allocation; online expectation-maximization; big data; big model; lifelong topic modeling

类别

Computer Science, Artificial Intelligence Computer Science, Information Systems Engineering, Electrical & Electronic

资金

NSFC [61373092, 61033013, 61272449, 61572339]
Natural Science Foundation of Jiangsu Higher Education Institutions of China [12KJA520004]
Innovative Research Team in Soochow University [SDT2012B02]
GRF grant from RGC UGC Hong Kong (GRF) [9041574]
City University of Hong Kong [7008026]

向作者/读者索取更多资源

Protocol

社区支持

Reagent

社区支持

摘要

The expectation-maximization (EM) algorithm can compute the maximum-likelihood (ML) or maximum a posterior (MAP) point estimate of the mixture models or latent variable models such as latent Dirichlet allocation (LDA), which has been one of the most popular probabilistic topic modeling methods in the past decade. However, batch EM has high time and space complexities to learn big LDA models from big data streams. In this paper, we present a fast online EM (FOEM) algorithm that infers the topic distribution from the previously unseen documents incrementally with constant memory requirements. Within the stochastic approximation framework, we show that FOEM can converge to the local stationary point of the LDA's likelihood function. By dynamic scheduling for the fast speed and parameter streaming for the low memory usage, FOEM is more efficient for some lifelong topic modeling tasks than the state-of-the-art online LDA algorithms to handle both big data and big models (aka, big topic modeling) on just a PC.

Fast Online EM for Big Topic Modeling

期刊

IEEE TRANSACTIONS ON KNOWLEDGE AND DATA ENGINEERING

出版社

IEEE COMPUTER SOC

关键词

类别

资金

向作者/读者索取更多资源

Protocol

Reagent

作者

我是这篇论文的作者

评论

主要评分

次要评分

新颖性

重要性

科学严谨性

评价这篇论文

推荐

Fast Online EM for Big Topic Modeling

期刊

IEEE TRANSACTIONS ON KNOWLEDGE AND DATA ENGINEERING

出版社

IEEE COMPUTER SOC

关键词

类别

资金

向作者/读者索取更多资源

Protocol

Reagent

作者

我是这篇论文的作者

评论

主要评分

次要评分

新颖性

重要性

科学严谨性

评价这篇论文

推荐

导出引文

分享论文