VinVL+L: Enriching Visual Representation with Location Context in VQA
The result's identifiers
Result code in IS VaVaI
<a href="https://www.isvavai.cz/riv?ss=detail&h=RIV%2F49777513%3A23520%2F23%3A43968165" target="_blank" >RIV/49777513:23520/23:43968165 - isvavai.cz</a>
Result on the web
<a href="https://ceur-ws.org/Vol-3349/paper4.pdf" target="_blank" >https://ceur-ws.org/Vol-3349/paper4.pdf</a>
DOI - Digital Object Identifier
—
Alternative languages
Result language
angličtina
Original language name
VinVL+L: Enriching Visual Representation with Location Context in VQA
Original language description
In this paper, we describe a novel method - VinVL+L - that enriches the visual representations (i.e. object tags and region features) of the State-of-the-Art Vision and Language (VL) method - VinVL - with Location information. To verify the importance of such metadata for VL models, we (i) trained a Swin-B model on the Places365 dataset and obtained additional sets of visual and tag features; both were made public to allow reproducibility and further experiments, (ii) did an architectural update to the existing VinVL method to include the new feature sets, and (iii) provide a qualitative and quantitative evaluation. By including just binary location metadata, the VinVL+L method provides incremental improvement to the State-of-the-Art VinVL in Visual Question Answering (VQA). The VinVL+L achieved an accuracy of 64.85% and increased the performance by +0.32% in terms of accuracy on the GQA dataset; the statistical significance of the new representations is verified via Approximate Randomization. The code and newly generated sets of features are available at https://github.com/vyskocj/VinVL-L.
Czech name
—
Czech description
—
Classification
Type
D - Article in proceedings
CEP classification
—
OECD FORD branch
20205 - Automation and control systems
Result continuities
Project
—
Continuities
S - Specificky vyzkum na vysokych skolach
Others
Publication year
2023
Confidentiality
S - Úplné a pravdivé údaje o projektu nepodléhají ochraně podle zvláštních právních předpisů
Data specific for result type
Article name in the collection
CEUR Workshop Proceedings
ISBN
—
ISSN
1613-0073
e-ISSN
—
Number of pages
9
Pages from-to
1-9
Publisher name
CEUR-WS
Place of publication
Aachen
Event location
Kremže, Rakousko
Event date
Feb 15, 2023
Type of event by nationality
EUR - Evropská akce
UT code for WoS article
—