Assigning scientific texts to existing ontologies
Identifikátory výsledku
Kód výsledku v IS VaVaI
<a href="https://www.isvavai.cz/riv?ss=detail&h=RIV%2F67985807%3A_____%2F25%3A00643769" target="_blank" >RIV/67985807:_____/25:00643769 - isvavai.cz</a>
Výsledek na webu
<a href="https://doi.org/10.15439/2025F1850" target="_blank" >https://doi.org/10.15439/2025F1850</a>
DOI - Digital Object Identifier
<a href="http://dx.doi.org/10.15439/2025F1850" target="_blank" >10.15439/2025F1850</a>
Alternativní jazyky
Jazyk výsledku
angličtina
Název v původním jazyce
Assigning scientific texts to existing ontologies
Popis výsledku v původním jazyce
Humans try to help computers understand the properties of the real world, and ontologies can be used for this task. Scientists publish their research in papers, and their results should be used to improve existing ontologies to be up-to-date. Manual enhancement of ontologies is highly time-consuming for domain experts. This paper proposes a solution to match a scientific text to the most relevant ontology using artificial neural networks. Our approach selects a paragraph or a sentence, uses representation learning to embed it into a vector space by some embedder, and measures its relevance to embedded textual properties from the selected ontology by a modified version of a Siamese neural network. A modification is based on the extension of one branch of the Siamese network to aggregate inputs from a group of embeddings. We have considered different embedders, in particular two variants of BERT, InferSent, GloVe with TF-IDF weighted mean, Doc2Vec in the distributed memory variant, and the Llama 3.1 with LLM2vec framework. Their quality has been evaluated on a use case with available ontologies from several application domains. The best results were achieved with InferSent and SentenceBERT.
Název v anglickém jazyce
Assigning scientific texts to existing ontologies
Popis výsledku anglicky
Humans try to help computers understand the properties of the real world, and ontologies can be used for this task. Scientists publish their research in papers, and their results should be used to improve existing ontologies to be up-to-date. Manual enhancement of ontologies is highly time-consuming for domain experts. This paper proposes a solution to match a scientific text to the most relevant ontology using artificial neural networks. Our approach selects a paragraph or a sentence, uses representation learning to embed it into a vector space by some embedder, and measures its relevance to embedded textual properties from the selected ontology by a modified version of a Siamese neural network. A modification is based on the extension of one branch of the Siamese network to aggregate inputs from a group of embeddings. We have considered different embedders, in particular two variants of BERT, InferSent, GloVe with TF-IDF weighted mean, Doc2Vec in the distributed memory variant, and the Llama 3.1 with LLM2vec framework. Their quality has been evaluated on a use case with available ontologies from several application domains. The best results were achieved with InferSent and SentenceBERT.
Klasifikace
Druh
D - Stať ve sborníku
CEP obor
—
OECD FORD obor
10201 - Computer sciences, information science, bioinformathics (hardware development to be 2.2, social aspect to be 5.8)
Návaznosti výsledku
Projekt
—
Návaznosti
I - Institucionalni podpora na dlouhodoby koncepcni rozvoj vyzkumne organizace
Ostatní
Rok uplatnění
2025
Kód důvěrnosti údajů
S - Úplné a pravdivé údaje o projektu nepodléhají ochraně podle zvláštních právních předpisů
Údaje specifické pro druh výsledku
Název statě ve sborníku
Proceedings of the 20th Conference on Computer Science and Intelligence Systems (FedCSIS)
ISBN
—
ISSN
2300-5963
e-ISSN
—
Počet stran výsledku
9
Strana od-do
185-193
Název nakladatele
IEEE
Místo vydání
Piscataway
Místo konání akce
Krakow
Datum konání akce
14. 9. 2025
Typ akce podle státní příslušnosti
WRD - Celosvětová akce
Kód UT WoS článku
—