4.2 Article

Latent semantic indexing: A probabilistic analysis

Journal

JOURNAL OF COMPUTER AND SYSTEM SCIENCES
Volume 61, Issue 2, Pages 217-235

Publisher

ACADEMIC PRESS INC ELSEVIER SCIENCE
DOI: 10.1006/jcss.2000.1711

Keywords

-

Ask authors/readers for more resources

Latent semantic indexing (LSI) is an information retrieval technique based on the spectral analysis of the term-document matrix, whose empirical success had heretofore been without rigorous prediction and explanation. We prove that, under certain conditions, LSI does succeed in capturing the underlying semantics of the corpus and achieves improved retrieval performance. We propose the technique of random projection as a way of speeding up LSI. We complement our theorems with encouraging experimental results. We also argue that our results may be viewed in a more general framework, as a theoretical basis for the use of spectral methods in a wider class of applications such as collaborative filtering. (C) 2000 Academic Press.

Authors

I am an author on this paper
Click your name to claim this paper and add it to your profile.

Reviews

Primary Rating

4.2
Not enough ratings

Secondary Ratings

Novelty
-
Significance
-
Scientific rigor
-
Rate this paper

Recommended

No Data Available
No Data Available