Beyond Image-Text Matching: Verb Understanding in Multimodal Transformers Using Guided Masking
Identifikátory výsledku
Kód výsledku v IS VaVaI
<a href="https://www.isvavai.cz/riv?ss=detail&h=RIV%2F00216305%3A26230%2F26%3A0199780" target="_blank" >RIV/00216305:26230/26:0199780 - isvavai.cz</a>
Výsledek na webu
<a href="http://dx.doi.org/10.1007/978-3-031-82670-2_7" target="_blank" >http://dx.doi.org/10.1007/978-3-031-82670-2_7</a>
DOI - Digital Object Identifier
<a href="http://dx.doi.org/10.1007/978-3-031-82670-2_7" target="_blank" >10.1007/978-3-031-82670-2_7</a>
Alternativní jazyky
Jazyk výsledku
angličtina
Název v původním jazyce
Beyond Image-Text Matching: Verb Understanding in Multimodal Transformers Using Guided Masking
Popis výsledku v původním jazyce
Probing methods are widely used to evaluate the multimodal representations of vision-language models (VLMs), with dominant approaches relying on zero-shot performance in image-text matching tasks. These methods typically assess models on curated datasets focusing on linguistic aspects such as counting, relations, or attributes. This work uses a complementary probing strategy called guided masking. This approach selectively masks different modalities and evaluates the model’s ability to predict the masked word. We specifically focus on probing verbs, as their comprehension is crucial for understanding actions and relationships in images, and it presents a more challenging task than subjects, objects, or attributes comprehension. Our analysis targets VLMs that use region-of-interest (ROI) features obtained from object detectors as input tokens. Our experiments demonstrate that selected models can accurately predict the correct verb, challenging previous conclusions based on image-text matching methods, which suggested VLMs fail in situations requiring verb understanding. The code for experiments will be available https://github.com/ivana-13/guided_masking.
Název v anglickém jazyce
Beyond Image-Text Matching: Verb Understanding in Multimodal Transformers Using Guided Masking
Popis výsledku anglicky
Probing methods are widely used to evaluate the multimodal representations of vision-language models (VLMs), with dominant approaches relying on zero-shot performance in image-text matching tasks. These methods typically assess models on curated datasets focusing on linguistic aspects such as counting, relations, or attributes. This work uses a complementary probing strategy called guided masking. This approach selectively masks different modalities and evaluates the model’s ability to predict the masked word. We specifically focus on probing verbs, as their comprehension is crucial for understanding actions and relationships in images, and it presents a more challenging task than subjects, objects, or attributes comprehension. Our analysis targets VLMs that use region-of-interest (ROI) features obtained from object detectors as input tokens. Our experiments demonstrate that selected models can accurately predict the correct verb, challenging previous conclusions based on image-text matching methods, which suggested VLMs fail in situations requiring verb understanding. The code for experiments will be available https://github.com/ivana-13/guided_masking.
Klasifikace
Druh
D - Stať ve sborníku
CEP obor
—
OECD FORD obor
10102 - Applied mathematics
Návaznosti výsledku
Projekt
—
Návaznosti
I - Institucionalni podpora na dlouhodoby koncepcni rozvoj vyzkumne organizace
Ostatní
Rok uplatnění
2025
Kód důvěrnosti údajů
S - Úplné a pravdivé údaje o projektu nepodléhají ochraně podle zvláštních právních předpisů
Údaje specifické pro druh výsledku
Název statě ve sborníku
SOFSEM 2025: Theory and Practice of Computer Science
ISBN
978-3-031-82669-6
ISSN
—
e-ISSN
1611-3349
Počet stran výsledku
14
Strana od-do
80-93
Název nakladatele
Springer Nature
Místo vydání
CHAM
Místo konání akce
Bratislava, Slovakia
Datum konání akce
20. 1. 2025
Typ akce podle státní příslušnosti
WRD - Celosvětová akce
Kód UT WoS článku
001534175600009