☆ 4.7 Article

SeqViews2SeqLabels: Learning 3D Global Features via Aggregating Sequential Views by RNN With Attention

IEEE TRANSACTIONS ON IMAGE PROCESSING (2019)

期刊

IEEE TRANSACTIONS ON IMAGE PROCESSING

卷 28, 期 2, 页码 658-672

出版社

IEEE-INST ELECTRICAL ELECTRONICS ENGINEERS INC

DOI: 10.1109/TIP.2018.2868426

关键词

3D feature learning; sequential views; sequential labels; view aggregation; RNN; attention

类别

Computer Science, Artificial Intelligence Engineering, Electrical & Electronic

资金

National Key RAMP
D Program of China [2018YFB0505400]
National Natural Science Foundation of China [61472202, 61672430]
Swiss National Science Foundation [169151, MYRG2018-00138-FST, MYRG2016-00134-FST, FDCT/273/2017/A]
NWPU Basic Research Fund [3102018jcc001]

向作者/读者索取更多资源

Protocol

社区支持

Reagent

社区支持

摘要

Learning 3D global features by aggregating multiple views has been introduced as a successful strategy for 3D shape analysis. In recent deep learning models with end-to-end training, pooling is a widely adopted procedure for view aggregation. However, pooling merely retains the max or mean value over all views, which disregards the content information of almost all views and also the spatial information among the views. To resolve these issues, we propose Sequential Views To Sequential Labels (SeqViews2SeqLabels) as a novel deep learning model with an encoder-decoder structure based on recurrent neural networks (RNNs) with attention. SeqViews2SeqLabels consists of two connected parts, an encoder-RNN followed by a decoder-RNN, that aim to learn the global features by aggregating sequential views and then performing shape classification from the learned global features, respectively. Specifically, the encoder-RNN learns the global features by simultaneously encoding the spatial and content information of sequential views, which captures the semantics of the view sequence. With the proposed prediction of sequential labels, the decoder-RNN performs more accurate classification using the learned global features by predicting sequential labels step by step. Learning to predict sequential labels provides more and finer discriminative information among shape classes to learn, which alleviates the overfitting problem inherent in training using a limited number of 3D shapes. Moreover, we introduce an attention mechanism to further improve the discriminative ability of SeqViews2SeqLabels. This mechanism increases the weight of views that are distinctive to each shape class, and it dramatically reduces the effect of selecting the first view position. Shape classification and retrieval results under three large-scale benchmarks verify that SeqViews2SeqLabels learns more discriminative global features by more effectively aggregating sequential views than state-of-the-art methods.

SeqViews2SeqLabels: Learning 3D Global Features via Aggregating Sequential Views by RNN With Attention

期刊

IEEE TRANSACTIONS ON IMAGE PROCESSING

出版社

IEEE-INST ELECTRICAL ELECTRONICS ENGINEERS INC

关键词

类别

资金

向作者/读者索取更多资源

Protocol

Reagent

作者

我是这篇论文的作者

评论

主要评分

次要评分

新颖性

重要性

科学严谨性

评价这篇论文

推荐

SeqViews2SeqLabels: Learning 3D Global Features via Aggregating Sequential Views by RNN With Attention

期刊

IEEE TRANSACTIONS ON IMAGE PROCESSING

出版社

IEEE-INST ELECTRICAL ELECTRONICS ENGINEERS INC

关键词

类别

资金

向作者/读者索取更多资源

Protocol

Reagent

作者

我是这篇论文的作者

评论

主要评分

次要评分

新颖性

重要性

科学严谨性

评价这篇论文

推荐

导出引文

分享论文