SynCSE: syntax graph-based contrastive learning of sentence embeddings
Identifikátory výsledku
Kód výsledku v IS VaVaI
<a href="https://www.isvavai.cz/riv?ss=detail&h=RIV%2F00216208%3A11320%2F26%3AF8REJIEU" target="_blank" >RIV/00216208:11320/26:F8REJIEU - isvavai.cz</a>
Výsledek na webu
<a href="http://dx.doi.org/10.1016/j.eswa.2025.128047" target="_blank" >http://dx.doi.org/10.1016/j.eswa.2025.128047</a>
DOI - Digital Object Identifier
<a href="http://dx.doi.org/10.1016/j.eswa.2025.128047" target="_blank" >10.1016/j.eswa.2025.128047</a>
Alternativní jazyky
Jazyk výsledku
angličtina
Název v původním jazyce
SynCSE: syntax graph-based contrastive learning of sentence embeddings
Popis výsledku v původním jazyce
Pre-trained language models (PrLMs) trained via contrastive learning methods achieved state-of-the-art performance on various natural language processing (NLP) tasks. Most PrLMs for sentence embedding focuses on context similarity as an objective function of contrastive learning. However, we found that these PrLMs, including recently released large language models (LLMs) like LLaMA, underperform when analyzing syntax information on probing tasks. This limitation becomes particularly noticeable in applications that depend on nuanced sentence understanding, such as the Retrieval Augmented Generation (RAG) framework in LLMs. This paper introduces a new sentence embedding model named SynCSE: Syntax Graph-based Contrastive Learning of Sentence Embeddings. Our approach enables meaningful sentence embeddings of language models through learning the syntactic features. To accomplish this, we train a PrLM with graph neural networks (GNNs) receiving a directed syntax graph. We then detach additional GNN layers from PrLM for inference; which does not require a syntax graph. The proposed model gains improvement on baselines in sentence textual similarity (STS) tasks, transfer tasks, and especially probing tasks. Additionally, we observe that our model has improved alignment and competitive uniformity compared to the baseline. © 2025 Elsevier Ltd
Název v anglickém jazyce
SynCSE: syntax graph-based contrastive learning of sentence embeddings
Popis výsledku anglicky
Pre-trained language models (PrLMs) trained via contrastive learning methods achieved state-of-the-art performance on various natural language processing (NLP) tasks. Most PrLMs for sentence embedding focuses on context similarity as an objective function of contrastive learning. However, we found that these PrLMs, including recently released large language models (LLMs) like LLaMA, underperform when analyzing syntax information on probing tasks. This limitation becomes particularly noticeable in applications that depend on nuanced sentence understanding, such as the Retrieval Augmented Generation (RAG) framework in LLMs. This paper introduces a new sentence embedding model named SynCSE: Syntax Graph-based Contrastive Learning of Sentence Embeddings. Our approach enables meaningful sentence embeddings of language models through learning the syntactic features. To accomplish this, we train a PrLM with graph neural networks (GNNs) receiving a directed syntax graph. We then detach additional GNN layers from PrLM for inference; which does not require a syntax graph. The proposed model gains improvement on baselines in sentence textual similarity (STS) tasks, transfer tasks, and especially probing tasks. Additionally, we observe that our model has improved alignment and competitive uniformity compared to the baseline. © 2025 Elsevier Ltd
Klasifikace
Druh
J<sub>SC</sub> - Článek v periodiku v databázi SCOPUS
CEP obor
—
OECD FORD obor
10201 - Computer sciences, information science, bioinformathics (hardware development to be 2.2, social aspect to be 5.8)
Návaznosti výsledku
Projekt
—
Návaznosti
—
Ostatní
Rok uplatnění
2025
Kód důvěrnosti údajů
S - Úplné a pravdivé údaje o projektu nepodléhají ochraně podle zvláštních právních předpisů
Údaje specifické pro druh výsledku
Název periodika
Expert Systems with Applications
ISSN
0957-4174
e-ISSN
—
Svazek periodika
287
Číslo periodika v rámci svazku
2025
Stát vydavatele periodika
US - Spojené státy americké
Počet stran výsledku
16
Strana od-do
128047
Kód UT WoS článku
—
EID výsledku v databázi Scopus
2-s2.0-105005497338