☆ 4.4 Article

Convolutional recurrent neural networks with hidden Markov model bootstrap for scene text recognition

IET COMPUTER VISION (2017)

期刊

IET COMPUTER VISION

卷 11, 期 6, 页码 497-504

出版社

WILEY

DOI: 10.1049/iet-cvi.2016.0417

关键词

recurrent neural nets; text detection; hidden Markov models; convolutional recurrent neural networks; scene text recognition; RNN; CNN; Gaussian mixture model-hidden Markov model; lexicon-free text; lexicon-based text

类别

Computer Science, Artificial Intelligence Engineering, Electrical & Electronic

资金

National Natural Science Foundation of China [6117019]

向作者/读者索取更多资源

Protocol

社区支持

Reagent

社区支持

摘要

Text recognition in natural scene remains a challenging problem due to the highly variable appearance in unconstrained condition. The authors develop a system that directly transcribes scene text images to text without character segmentation. They formulate the problem as sequence labelling. They build a convolutional recurrent neural network (RNN) by using deep convolutional neural networks (CNN) for modelling text appearance and RNNs for sequence dynamics. The two models are complementary in modelling capabilities and so integrated together to form the segmentation free system. They train a Gaussian mixture model-hidden Markov model to supervise the training of the CNN model. The system is data driven and needs no hand labelled training data. Their method has several appealing properties: (i) It can recognise arbitrary length text images. (ii) The recognition process does not involve sophisticated character segmentation. (iii) It is trained on scene text images with only word-level transcriptions. (iv) It can recognise both the lexicon-based or lexicon-free text. The proposed system achieves competitive performance comparison with the state of the art on several public scene text datasets, including both lexicon-based and non-lexicon ones.

Convolutional recurrent neural networks with hidden Markov model bootstrap for scene text recognition

期刊

IET COMPUTER VISION

出版社

WILEY

关键词

类别

资金

向作者/读者索取更多资源

Protocol

Reagent

作者

我是这篇论文的作者

评论

主要评分

次要评分

新颖性

重要性

科学严谨性

评价这篇论文

推荐

Convolutional recurrent neural networks with hidden Markov model bootstrap for scene text recognition

期刊

IET COMPUTER VISION

出版社

WILEY

关键词

类别

资金

向作者/读者索取更多资源

Protocol

Reagent

作者

我是这篇论文的作者

评论

主要评分

次要评分

新颖性

重要性

科学严谨性

评价这篇论文

推荐

导出引文

分享论文