Experimental Evaluation of Static Image Sub-Region-Based Search Models Using CLIP
Identifikátory výsledku
Kód výsledku v IS VaVaI
<a href="https://www.isvavai.cz/riv?ss=detail&h=RIV%2F00216208%3A11320%2F25%3A10504478" target="_blank" >RIV/00216208:11320/25:10504478 - isvavai.cz</a>
Výsledek na webu
<a href="http://dx.doi.org/10.1007/978-3-032-06069-3_12" target="_blank" >http://dx.doi.org/10.1007/978-3-032-06069-3_12</a>
DOI - Digital Object Identifier
<a href="http://dx.doi.org/10.1007/978-3-032-06069-3_12" target="_blank" >10.1007/978-3-032-06069-3_12</a>
Alternativní jazyky
Jazyk výsledku
angličtina
Název v původním jazyce
Experimental Evaluation of Static Image Sub-Region-Based Search Models Using CLIP
Popis výsledku v původním jazyce
Advances in multimodal text-image models have enabled effective text-based querying in extensive image collections. While these models show convincing performance for everyday life scenes, querying in highly homogeneous, specialized domains remains challenging. The primary problem is that users can often provide only vague textual descriptions as they lack expert knowledge to discriminate between homogenous entities. This work investigates whether adding location-based prompts to complement these vague text queries can enhance retrieval performance. Specifically, we collected a dataset of 741 human annotations, each containing short and long textual descriptions and bounding boxes indicating regions of interest in challenging underwater scenes. Using these annotations, we evaluate the performance of CLIP when queried on various static sub-regions of images compared to the full image. Our results show that both a simple 3-by-3 partitioning and a 5-grid overlap significantly improve retrieval effectiveness and remain robust to perturbations of the annotation box.
Název v anglickém jazyce
Experimental Evaluation of Static Image Sub-Region-Based Search Models Using CLIP
Popis výsledku anglicky
Advances in multimodal text-image models have enabled effective text-based querying in extensive image collections. While these models show convincing performance for everyday life scenes, querying in highly homogeneous, specialized domains remains challenging. The primary problem is that users can often provide only vague textual descriptions as they lack expert knowledge to discriminate between homogenous entities. This work investigates whether adding location-based prompts to complement these vague text queries can enhance retrieval performance. Specifically, we collected a dataset of 741 human annotations, each containing short and long textual descriptions and bounding boxes indicating regions of interest in challenging underwater scenes. Using these annotations, we evaluate the performance of CLIP when queried on various static sub-regions of images compared to the full image. Our results show that both a simple 3-by-3 partitioning and a 5-grid overlap significantly improve retrieval effectiveness and remain robust to perturbations of the annotation box.
Klasifikace
Druh
D - Stať ve sborníku
CEP obor
—
OECD FORD obor
10201 - Computer sciences, information science, bioinformathics (hardware development to be 2.2, social aspect to be 5.8)
Návaznosti výsledku
Projekt
<a href="/cs/project/GA25-16785S" target="_blank" >GA25-16785S: Využití Large Language Modelů v Multi-Objective Doporučovacích Systémech</a><br>
Návaznosti
P - Projekt vyzkumu a vyvoje financovany z verejnych zdroju (s odkazem do CEP)
Ostatní
Rok uplatnění
2025
Kód důvěrnosti údajů
S - Úplné a pravdivé údaje o projektu nepodléhají ochraně podle zvláštních právních předpisů
Údaje specifické pro druh výsledku
Název statě ve sborníku
18th International Conference, SISAP 2025
ISBN
978-3-032-06069-3
ISSN
—
e-ISSN
—
Počet stran výsledku
14
Strana od-do
140-153
Název nakladatele
Springer Nature Switzerland AG
Místo vydání
Cham
Místo konání akce
Reykjavik, Iceland
Datum konání akce
1. 10. 2025
Typ akce podle státní příslušnosti
WRD - Celosvětová akce
Kód UT WoS článku
—