Visual Question Answering (VQA) increasingly attracts industry and academia attention. It requires the model to provide a natural language answer by an image and a related natural language question. Meanwhile, it relates to multidisciplinary research such as natural language understanding, visual information retrieval, and multimodal reasoning. As a multimodality task,...