Less Is More: Similarity Models for Content-Based Video Retrieval
The result's identifiers
Result code in IS VaVaI
<a href="https://www.isvavai.cz/riv?ss=detail&h=RIV%2F00216208%3A11320%2F23%3A10468869" target="_blank" >RIV/00216208:11320/23:10468869 - isvavai.cz</a>
Result on the web
<a href="https://doi.org/10.1007/978-3-031-27818-1_5" target="_blank" >https://doi.org/10.1007/978-3-031-27818-1_5</a>
DOI - Digital Object Identifier
<a href="http://dx.doi.org/10.1007/978-3-031-27818-1_5" target="_blank" >10.1007/978-3-031-27818-1_5</a>
Alternative languages
Result language
angličtina
Original language name
Less Is More: Similarity Models for Content-Based Video Retrieval
Original language description
The concept of object-to-object similarity plays a crucial role in interactive content-based video retrieval tools. Similarity (or distance) models are core components of several retrieval concepts, e.g. Query by Example or relevance feedback. In these scenarios, the common approach is to apply some feature extractor that transforms the object to a vector of features, i.e., positions it into an induced latent space. The similarity is then based on some distance metric in this space. Historically, feature extractors were mostly based on some color histograms or hand-crafted descriptors such as SIFT, but nowadays state-of-the-art tools mostly rely on some deep learning (DL) approaches. However, so far there were no systematic study of how suitable are individual feature extractors in the video retrieval domain. Or, in other words, to what extent are human-perceived and model-based similarities concordant. To fill this gap, we conducted a user study with over 4000 similarity judgements comparing over 20 variants of feature extractors. Results corroborate the dominance of deep learning approaches, but surprisingly favor smaller and simpler DL models instead of larger ones.
Czech name
—
Czech description
—
Classification
Type
D - Article in proceedings
CEP classification
—
OECD FORD branch
10201 - Computer sciences, information science, bioinformathics (hardware development to be 2.2, social aspect to be 5.8)
Result continuities
Project
<a href="/en/project/GA22-21696S" target="_blank" >GA22-21696S: Deep Visual Representations of Unstructured Data</a><br>
Continuities
P - Projekt vyzkumu a vyvoje financovany z verejnych zdroju (s odkazem do CEP)
Others
Publication year
2023
Confidentiality
S - Úplné a pravdivé údaje o projektu nepodléhají ochraně podle zvláštních právních předpisů
Data specific for result type
Article name in the collection
MULTIMEDIA MODELING, MMM 2023, PT II
ISBN
978-3-031-27817-4
ISSN
0302-9743
e-ISSN
1611-3349
Number of pages
12
Pages from-to
54-65
Publisher name
SPRINGER INTERNATIONAL PUBLISHING AG
Place of publication
CHAM
Event location
Bergen
Event date
Jan 9, 2023
Type of event by nationality
WRD - Celosvětová akce
UT code for WoS article
000996578000005