☆ 4.6 Article

Multi-Modality Global Fusion Attention Network for Visual Question Answering

ELECTRONICS (2020)

Journal

ELECTRONICS

Volume 9, Issue 11, Pages -

Publisher

MDPI

DOI: 10.3390/electronics9111882

Keywords

visual question answering; global attention mechanism; deep learning

Funding

National Key Research and Development Program of China [2019YFC0118200]

Ask authors/readers for more resources

Protocol

Community support

Reagent

Community support

Abstract

Visual question answering (VQA) requires a high-level understanding of both questions and images, along with visual reasoning to predict the correct answer. Therefore, it is important to design an effective attention model to associate key regions in an image with key words in a question. Up to now, most attention-based approaches only model the relationships between individual regions in an image and words in a question. It is not enough to predict the correct answer for VQA, as human beings always think in terms of global information, not only local information. In this paper, we propose a novel multi-modality global fusion attention network (MGFAN) consisting of stacked global fusion attention (GFA) blocks, which can capture information from global perspectives. Our proposed method computes co-attention and self-attention at the same time, rather than computing them individually. We validate our proposed method on the two most commonly used benchmarks, the VQA-v2 datasets. Experimental results show that the proposed method outperforms the previous state-of-the-art. Our best single model achieves 70.67% accuracy on the test-dev set of VQA-v2.

Multi-Modality Global Fusion Attention Network for Visual Question Answering

Journal

ELECTRONICS

Publisher

MDPI

Keywords

Categories

Funding

Ask authors/readers for more resources

Protocol

Reagent

Authors

I am an author on this paper

Reviews

Primary Rating

Secondary Ratings

Novelty

Significance

Scientific rigor

Rate this paper

Recommended

Multi-Modality Global Fusion Attention Network for Visual Question Answering

Journal

ELECTRONICS

Publisher

MDPI

Keywords

Categories

Funding

Ask authors/readers for more resources

Protocol

Reagent

Authors

I am an author on this paper

Reviews

Primary Rating

Secondary Ratings

Novelty

Significance

Scientific rigor

Rate this paper

Recommended

Export Citation

Share Paper