☆ 4.7 Article

Scene-Text Oriented Referring Expression Comprehension

IEEE TRANSACTIONS ON MULTIMEDIA (2023)

期刊

IEEE TRANSACTIONS ON MULTIMEDIA

卷 25, 期 -, 页码 7208-7221

出版社

IEEE-INST ELECTRICAL ELECTRONICS ENGINEERS INC

DOI: 10.1109/TMM.2022.3219642

关键词

Visualization; Task analysis; Semantics; Text recognition; Pain; Natural languages; Image recognition; Referring expression comprehension; scene text representation; multimodal alignment

类别

Computer Science, Information Systems Computer Science, Software Engineering Telecommunications

向作者/读者索取更多资源

Protocol

社区支持

Reagent

社区支持

智能总结 New
摘要

This article introduces a new task called scene-text oriented referring expression comprehension and proposes a scene text awareness network to address alignment and error issues. Experimental results show that the proposed method effectively comprehends scene-text oriented referring expressions and achieves excellent performance.

Referring expression comprehension (REC) aims to identify and locate a specific object in visual scenes referred to by a natural language expression. Existing studies of REC only focus on basic visual attributes and neglect scene text. Since scene text has the functions of object identification and disambiguation, it is naturally and frequently used to refer to objects. However, existing methods do not explicitly recognize text in images and fail to align scene text mentioned in expressions with the text shown in images, resulting in object localization errors. This article takes the first step toward addressing these limitations. First, we introduce a new task called scene-text oriented referring expression comprehension, which aims to align visual cues and textual semantics of scene text with referring expressions and visual contents. Second, we propose a scene text awareness network that can bridge the gap between texts from two modalities by grounding visual representations of expression-correlated scene texts. Specifically, we propose a correlated text extraction module to solve the problem of lacking semantic understanding, and a correlated region activation module to address the fixed alignment problem and absent alignment problem. These modules ensure that the proposed method focuses on local regions that are most relevant to scene text, thus mitigating the misalignment of scene text with irrelevant regions. Third, to conduct quantitative evaluations, we establish a new benchmark dataset called RefText. Experimental results demonstrate that the proposed method can effectively comprehend scene-text oriented referring expressions and achieves excellent performance.

Scene-Text Oriented Referring Expression Comprehension

期刊

IEEE TRANSACTIONS ON MULTIMEDIA

出版社

IEEE-INST ELECTRICAL ELECTRONICS ENGINEERS INC

关键词

类别

向作者/读者索取更多资源

Protocol

Reagent

作者

我是这篇论文的作者

评论

主要评分

次要评分

新颖性

重要性

科学严谨性

评价这篇论文

推荐

Scene-Text Oriented Referring Expression Comprehension

期刊

IEEE TRANSACTIONS ON MULTIMEDIA

出版社

IEEE-INST ELECTRICAL ELECTRONICS ENGINEERS INC

关键词

类别

向作者/读者索取更多资源

Protocol

Reagent

作者

我是这篇论文的作者

评论

主要评分

次要评分

新颖性

重要性

科学严谨性

评价这篇论文

推荐

导出引文

分享论文