Towards Personalized Similarity Search for Vector Databases
Identifikátory výsledku
Kód výsledku v IS VaVaI
<a href="https://www.isvavai.cz/riv?ss=detail&h=RIV%2F00216224%3A14330%2F25%3A00140292" target="_blank" >RIV/00216224:14330/25:00140292 - isvavai.cz</a>
Výsledek na webu
<a href="http://dx.doi.org/10.1007/978-3-031-75823-2_11" target="_blank" >http://dx.doi.org/10.1007/978-3-031-75823-2_11</a>
DOI - Digital Object Identifier
<a href="http://dx.doi.org/10.1007/978-3-031-75823-2_11" target="_blank" >10.1007/978-3-031-75823-2_11</a>
Alternativní jazyky
Jazyk výsledku
angličtina
Název v původním jazyce
Towards Personalized Similarity Search for Vector Databases
Popis výsledku v původním jazyce
The importance of similarity search has become prominent in the fast-evolving vector databases, which apply content embedding techniques on complex data to produce and manage large collections of high-dimensional vectors. Processing of such data is only possible by using a similarity function for storage, structure, and retrieval. However, if multiple users access the collection, their views on similarity can differ as similarity, in general, is subjective and context-dependent. In this article, we elaborate on the problem of a similarity search engine implementation, where users use a common index but search with personalised views of similarity, implemented by a possibly different similarity model. Specifically, we define a foundational theoretical framework and conduct experiments on real-life data to confirm the viability of such an approach. The experiments also indicate future research directions needed to propose and implement an effective and efficient personalised similarity search engine.
Název v anglickém jazyce
Towards Personalized Similarity Search for Vector Databases
Popis výsledku anglicky
The importance of similarity search has become prominent in the fast-evolving vector databases, which apply content embedding techniques on complex data to produce and manage large collections of high-dimensional vectors. Processing of such data is only possible by using a similarity function for storage, structure, and retrieval. However, if multiple users access the collection, their views on similarity can differ as similarity, in general, is subjective and context-dependent. In this article, we elaborate on the problem of a similarity search engine implementation, where users use a common index but search with personalised views of similarity, implemented by a possibly different similarity model. Specifically, we define a foundational theoretical framework and conduct experiments on real-life data to confirm the viability of such an approach. The experiments also indicate future research directions needed to propose and implement an effective and efficient personalised similarity search engine.
Klasifikace
Druh
D - Stať ve sborníku
CEP obor
—
OECD FORD obor
10200 - Computer and information sciences
Návaznosti výsledku
Projekt
<a href="/cs/project/VK01010147" target="_blank" >VK01010147: Automatizovaná forenzní laboratoř digitálních dat pro odhalování komplexní trestné činnosti</a><br>
Návaznosti
P - Projekt vyzkumu a vyvoje financovany z verejnych zdroju (s odkazem do CEP)
Ostatní
Rok uplatnění
2025
Kód důvěrnosti údajů
S - Úplné a pravdivé údaje o projektu nepodléhají ochraně podle zvláštních právních předpisů
Údaje specifické pro druh výsledku
Název statě ve sborníku
17th International Conference on Similarity Search and Applications (SISAP 2024)
ISBN
9783031758225
ISSN
0302-9743
e-ISSN
1611-3349
Počet stran výsledku
14
Strana od-do
126-139
Název nakladatele
Springer
Místo vydání
Cham
Místo konání akce
Providence, RI, USA
Datum konání akce
1. 1. 2024
Typ akce podle státní příslušnosti
WRD - Celosvětová akce
Kód UT WoS článku
001422992900011