☆ 4.7 Article

Locally Supervised Deep Hybrid Model for Scene Recognition

IEEE TRANSACTIONS ON IMAGE PROCESSING (2017)

期刊

IEEE TRANSACTIONS ON IMAGE PROCESSING

卷 26, 期 2, 页码 808-820

出版社

IEEE-INST ELECTRICAL ELECTRONICS ENGINEERS INC

DOI: 10.1109/TIP.2016.2629443

关键词

Scene recognition; convolutional neural networks; local convolutional supervision; Fisher convolutional vctor

类别

Computer Science, Artificial Intelligence Engineering, Electrical & Electronic

资金

National High-Tech Research and Development Program of China [2016YFC1400704]
National Natural Science Foundation of China [61503367]
Guangdong Research Program [2014B050505017, 2015B010129013, 2015A030310289]
External Cooperation Program of BIC Chinese Academy of Sciences [172644KYSB20150019]
Shenzhen Research Program [JSGG20150925164740726, JCYJ20150925 163005055, CXZZ20150930104115529]

向作者/读者索取更多资源

Protocol

社区支持

Reagent

社区支持

摘要

Convolutional neural networks (CNNs) have recently achieved remarkable successes in various image classification and understanding tasks. The deep features obtained at the top fully connected layer of the CNN (FC-features) exhibit rich global semantic information and are extremely effective in image classification. On the other hand, the convolutional features in the middle layers of the CNN also contain meaningful local information, but are not fully explored for image representation. In this paper, we propose a novel locally supervised deep hybrid model (LS-DHM) that effectively enhances and explores the convolutional features for scene recognition. First, we notice that the convolutional features capture local objects and fine structures of scene images, which yield important cues for discriminating ambiguous scenes, whereas these features are significantly eliminated in the highly compressed FC representation. Second, we propose a new local convolutional supervision layer to enhance the local structure of the image by directly propagating the label information to the convolutional layers. Third, we propose an efficient Fisher convolutional vector (FCV) that successfully rescues the orderless mid-level semantic information (e.g., objects and textures) of scene image. The FCV encodes the large-sized convolutional maps into a fixed-length mid-level representation, and is demonstrated to be strongly complementary to the high-level FC-features. Finally, both the FCV and FC-features are collaboratively employed in the LS-DHM representation, which achieves outstanding performance in our experiments. It obtains 83.75% and 67.56% accuracies, respectively, on the heavily benchmarked MIT Indoor67 and SUN397 data sets, advancing the state-of-the-art substantially.

Locally Supervised Deep Hybrid Model for Scene Recognition

期刊

IEEE TRANSACTIONS ON IMAGE PROCESSING

出版社

IEEE-INST ELECTRICAL ELECTRONICS ENGINEERS INC

关键词

类别

资金

向作者/读者索取更多资源

Protocol

Reagent

作者

我是这篇论文的作者

评论

主要评分

次要评分

新颖性

重要性

科学严谨性

评价这篇论文

推荐

Locally Supervised Deep Hybrid Model for Scene Recognition

期刊

IEEE TRANSACTIONS ON IMAGE PROCESSING

出版社

IEEE-INST ELECTRICAL ELECTRONICS ENGINEERS INC

关键词

类别

资金

向作者/读者索取更多资源

Protocol

Reagent

作者

我是这篇论文的作者

评论

主要评分

次要评分

新颖性

重要性

科学严谨性

评价这篇论文

推荐

导出引文

分享论文