GDReCo: Fine-grained gene-disease relationship extraction corpus
Identifikátory výsledku
Kód výsledku v IS VaVaI
<a href="https://www.isvavai.cz/riv?ss=detail&h=RIV%2F00216208%3A11320%2F26%3AF4LJWVMB" target="_blank" >RIV/00216208:11320/26:F4LJWVMB - isvavai.cz</a>
Výsledek na webu
<a href="http://dx.doi.org/10.1016/j.cmpb.2025.108773" target="_blank" >http://dx.doi.org/10.1016/j.cmpb.2025.108773</a>
DOI - Digital Object Identifier
<a href="http://dx.doi.org/10.1016/j.cmpb.2025.108773" target="_blank" >10.1016/j.cmpb.2025.108773</a>
Alternativní jazyky
Jazyk výsledku
angličtina
Název v původním jazyce
GDReCo: Fine-grained gene-disease relationship extraction corpus
Popis výsledku v původním jazyce
Background and objective: Understanding gene-disease relationships is crucial for medical research, drug discovery, clinical diagnosis, and other fields. However, there is currently no high-quality, fine-grained corpus available for training Natural Language Processing (NLP) models, which have proven to be effective in knowledge extraction. Methods: This study introduces a novel ontology framework for gene-disease associations, addressing the absence of a formal descriptive system and training corpus for NLP models. Results: We developed the Gene Disease Relationship Extraction Corpus (GDReCo), a refined dataset of over 24,000+ cases, including 2300+ manually annotated and 22,000+ model-predicted instances. BERT-based models trained on this data achieved high F1-scores for ""event"" and ""rel"" relationships, validating its effectiveness for Gene-Disease Relationship Extraction (GDRE) tasks. Conclusions: GDReCo serves as a valuable resource for biomedical research, though ChatGPT's limitations in fine-grained relation extraction are noted. © 2025
Název v anglickém jazyce
GDReCo: Fine-grained gene-disease relationship extraction corpus
Popis výsledku anglicky
Background and objective: Understanding gene-disease relationships is crucial for medical research, drug discovery, clinical diagnosis, and other fields. However, there is currently no high-quality, fine-grained corpus available for training Natural Language Processing (NLP) models, which have proven to be effective in knowledge extraction. Methods: This study introduces a novel ontology framework for gene-disease associations, addressing the absence of a formal descriptive system and training corpus for NLP models. Results: We developed the Gene Disease Relationship Extraction Corpus (GDReCo), a refined dataset of over 24,000+ cases, including 2300+ manually annotated and 22,000+ model-predicted instances. BERT-based models trained on this data achieved high F1-scores for ""event"" and ""rel"" relationships, validating its effectiveness for Gene-Disease Relationship Extraction (GDRE) tasks. Conclusions: GDReCo serves as a valuable resource for biomedical research, though ChatGPT's limitations in fine-grained relation extraction are noted. © 2025
Klasifikace
Druh
J<sub>SC</sub> - Článek v periodiku v databázi SCOPUS
CEP obor
—
OECD FORD obor
10201 - Computer sciences, information science, bioinformathics (hardware development to be 2.2, social aspect to be 5.8)
Návaznosti výsledku
Projekt
—
Návaznosti
—
Ostatní
Rok uplatnění
2025
Kód důvěrnosti údajů
S - Úplné a pravdivé údaje o projektu nepodléhají ochraně podle zvláštních právních předpisů
Údaje specifické pro druh výsledku
Název periodika
Computer Methods and Programs in Biomedicine
ISSN
0169-2607
e-ISSN
—
Svazek periodika
266
Číslo periodika v rámci svazku
2025
Stát vydavatele periodika
US - Spojené státy americké
Počet stran výsledku
25
Strana od-do
108773
Kód UT WoS článku
—
EID výsledku v databázi Scopus
2-s2.0-105002635469