Vše

Co hledáte?

Vše
Projekty
Výsledky výzkumu
Subjekty

Rychlé hledání

  • Projekty podpořené TA ČR
  • Významné projekty
  • Projekty s nejvyšší státní podporou
  • Aktuálně běžící projekty

Chytré vyhledávání

  • Takto najdu konkrétní +slovo
  • Takto z výsledků -slovo zcela vynechám
  • “Takto můžu najít celou frázi”

Beyond Image-Text Matching: Verb Understanding in Multimodal Transformers Using Guided Masking

Identifikátory výsledku

  • Kód výsledku v IS VaVaI

    <a href="https://www.isvavai.cz/riv?ss=detail&h=RIV%2F00216305%3A26230%2F26%3A0199780" target="_blank" >RIV/00216305:26230/26:0199780 - isvavai.cz</a>

  • Výsledek na webu

    <a href="http://dx.doi.org/10.1007/978-3-031-82670-2_7" target="_blank" >http://dx.doi.org/10.1007/978-3-031-82670-2_7</a>

  • DOI - Digital Object Identifier

    <a href="http://dx.doi.org/10.1007/978-3-031-82670-2_7" target="_blank" >10.1007/978-3-031-82670-2_7</a>

Alternativní jazyky

  • Jazyk výsledku

    angličtina

  • Název v původním jazyce

    Beyond Image-Text Matching: Verb Understanding in Multimodal Transformers Using Guided Masking

  • Popis výsledku v původním jazyce

    Probing methods are widely used to evaluate the multimodal representations of vision-language models (VLMs), with dominant approaches relying on zero-shot performance in image-text matching tasks. These methods typically assess models on curated datasets focusing on linguistic aspects such as counting, relations, or attributes. This work uses a complementary probing strategy called guided masking. This approach selectively masks different modalities and evaluates the model’s ability to predict the masked word. We specifically focus on probing verbs, as their comprehension is crucial for understanding actions and relationships in images, and it presents a more challenging task than subjects, objects, or attributes comprehension. Our analysis targets VLMs that use region-of-interest (ROI) features obtained from object detectors as input tokens. Our experiments demonstrate that selected models can accurately predict the correct verb, challenging previous conclusions based on image-text matching methods, which suggested VLMs fail in situations requiring verb understanding. The code for experiments will be available https://github.com/ivana-13/guided_masking.

  • Název v anglickém jazyce

    Beyond Image-Text Matching: Verb Understanding in Multimodal Transformers Using Guided Masking

  • Popis výsledku anglicky

    Probing methods are widely used to evaluate the multimodal representations of vision-language models (VLMs), with dominant approaches relying on zero-shot performance in image-text matching tasks. These methods typically assess models on curated datasets focusing on linguistic aspects such as counting, relations, or attributes. This work uses a complementary probing strategy called guided masking. This approach selectively masks different modalities and evaluates the model’s ability to predict the masked word. We specifically focus on probing verbs, as their comprehension is crucial for understanding actions and relationships in images, and it presents a more challenging task than subjects, objects, or attributes comprehension. Our analysis targets VLMs that use region-of-interest (ROI) features obtained from object detectors as input tokens. Our experiments demonstrate that selected models can accurately predict the correct verb, challenging previous conclusions based on image-text matching methods, which suggested VLMs fail in situations requiring verb understanding. The code for experiments will be available https://github.com/ivana-13/guided_masking.

Klasifikace

  • Druh

    D - Stať ve sborníku

  • CEP obor

  • OECD FORD obor

    10102 - Applied mathematics

Návaznosti výsledku

  • Projekt

  • Návaznosti

    I - Institucionalni podpora na dlouhodoby koncepcni rozvoj vyzkumne organizace

Ostatní

  • Rok uplatnění

    2025

  • Kód důvěrnosti údajů

    S - Úplné a pravdivé údaje o projektu nepodléhají ochraně podle zvláštních právních předpisů

Údaje specifické pro druh výsledku

  • Název statě ve sborníku

    SOFSEM 2025: Theory and Practice of Computer Science

  • ISBN

    978-3-031-82669-6

  • ISSN

  • e-ISSN

    1611-3349

  • Počet stran výsledku

    14

  • Strana od-do

    80-93

  • Název nakladatele

    Springer Nature

  • Místo vydání

    CHAM

  • Místo konání akce

    Bratislava, Slovakia

  • Datum konání akce

    20. 1. 2025

  • Typ akce podle státní příslušnosti

    WRD - Celosvětová akce

  • Kód UT WoS článku

    001534175600009