☆ 4.4 Article Proceedings Paper

Effective deep learning-based multi-modal retrieval

VLDB JOURNAL (2016)

期刊

VLDB JOURNAL

卷 25, 期 1, 页码 79-101

出版社

SPRINGER

DOI: 10.1007/s00778-015-0391-4

关键词

Deep learning; Multi-modal retrieval; Hashing; Auto-encoders; Deep convolutional neural network; Neural language model

类别

Computer Science, Hardware & Architecture Computer Science, Information Systems

资金

A*STAR Project [1321202073]
Human-Centered Cyber-physical Systems (HCCS) programme by A*STAR in Singapore

向作者/读者索取更多资源

Protocol

社区支持

Reagent

社区支持

摘要

Multi-modal retrieval is emerging as a new search paradigm that enables seamless information retrieval from various types of media. For example, users can simply snap a movie poster to search for relevant reviews and trailers. The mainstream solution to the problem is to learn a set of mapping functions that project data from different modalities into a common metric space in which conventional indexing schemes for high-dimensional space can be applied. Since the effectiveness of the mapping functions plays an essential role in improving search quality, in this paper, we exploit deep learning techniques to learn effective mapping functions. In particular, we first propose a general learning objective that effectively captures both intramodal and intermodal semantic relationships of data from heterogeneous sources. Given the general objective, we propose two learning algorithms to realize it: (1) an unsupervised approach that uses stacked auto-encoders and requires minimum prior knowledge on the training data and (2) a supervised approach using deep convolutional neural network and neural language model. Our training algorithms are memory efficient with respect to the data volume. Given a large training dataset, we split it into mini-batches and adjust the mapping functions continuously for each batch. Experimental results on three real datasets demonstrate that our proposed methods achieve significant improvement in search accuracy over the state-of-the-art solutions.

Effective deep learning-based multi-modal retrieval

期刊

VLDB JOURNAL

出版社

SPRINGER

关键词

类别

资金

向作者/读者索取更多资源

Protocol

Reagent

作者

我是这篇论文的作者

评论

主要评分

次要评分

新颖性

重要性

科学严谨性

评价这篇论文

推荐

Effective deep learning-based multi-modal retrieval

期刊

VLDB JOURNAL

出版社

SPRINGER

关键词

类别

资金

向作者/读者索取更多资源

Protocol

Reagent

作者

我是这篇论文的作者

评论

主要评分

次要评分

新颖性

重要性

科学严谨性

评价这篇论文

推荐

导出引文

分享论文