Leveraging large language models for literature-driven prioritization of protein binding pockets
Identifikátory výsledku
Kód výsledku v IS VaVaI
<a href="https://www.isvavai.cz/riv?ss=detail&h=RIV%2F61989592%3A15310%2F25%3A73635235" target="_blank" >RIV/61989592:15310/25:73635235 - isvavai.cz</a>
Výsledek na webu
<a href="https://academic.oup.com/bioinformatics/article/41/8/btaf449/8225722" target="_blank" >https://academic.oup.com/bioinformatics/article/41/8/btaf449/8225722</a>
DOI - Digital Object Identifier
<a href="http://dx.doi.org/10.1093/bioinformatics/btaf449" target="_blank" >10.1093/bioinformatics/btaf449</a>
Alternativní jazyky
Jazyk výsledku
angličtina
Název v původním jazyce
Leveraging large language models for literature-driven prioritization of protein binding pockets
Popis výsledku v původním jazyce
Motivation Accurately identifying and prioritizing protein binding pockets is a foundational element of small-molecule drug discovery. Defining these known pockets currently relies on a laborious manual process of extracting key residue data from selected publications, reconciling inconsistent terminology, and independently computing volumetric representations. This manual curation to ensure biological relevance is time-consuming, error-prone, and represents a major bottleneck for efficient, high-throughput drug discovery. Results We present a novel approach for the identification and prioritization of protein binding pockets for small molecules by combining geometric pocket detection with large language models (LLMs). Our method leverages Fpocket to generate candidate pockets, which are then validated against published experimental data extracted from research articles using LLM with a series of prompts fine-tuned to identify and extract residue-level information associated with experimentally confirmed binding sites. We developed a curated benchmark dataset of diverse proteins and associated literature to train and evaluate the LLM's performance in paper relevance assessment and pocket extraction. Availability and implementation The developed benchmark dataset and methodology are freely available at the GitHub repository (https://github.com/receptor-ai/LLM-benchmark-dataset) and Zenodo (DOI: 10.5281/zenodo.15798647).
Název v anglickém jazyce
Leveraging large language models for literature-driven prioritization of protein binding pockets
Popis výsledku anglicky
Motivation Accurately identifying and prioritizing protein binding pockets is a foundational element of small-molecule drug discovery. Defining these known pockets currently relies on a laborious manual process of extracting key residue data from selected publications, reconciling inconsistent terminology, and independently computing volumetric representations. This manual curation to ensure biological relevance is time-consuming, error-prone, and represents a major bottleneck for efficient, high-throughput drug discovery. Results We present a novel approach for the identification and prioritization of protein binding pockets for small molecules by combining geometric pocket detection with large language models (LLMs). Our method leverages Fpocket to generate candidate pockets, which are then validated against published experimental data extracted from research articles using LLM with a series of prompts fine-tuned to identify and extract residue-level information associated with experimentally confirmed binding sites. We developed a curated benchmark dataset of diverse proteins and associated literature to train and evaluate the LLM's performance in paper relevance assessment and pocket extraction. Availability and implementation The developed benchmark dataset and methodology are freely available at the GitHub repository (https://github.com/receptor-ai/LLM-benchmark-dataset) and Zenodo (DOI: 10.5281/zenodo.15798647).
Klasifikace
Druh
J<sub>imp</sub> - Článek v periodiku v databázi Web of Science
CEP obor
—
OECD FORD obor
10403 - Physical chemistry
Návaznosti výsledku
Projekt
—
Návaznosti
I - Institucionalni podpora na dlouhodoby koncepcni rozvoj vyzkumne organizace
Ostatní
Rok uplatnění
2025
Kód důvěrnosti údajů
S - Úplné a pravdivé údaje o projektu nepodléhají ochraně podle zvláštních právních předpisů
Údaje specifické pro druh výsledku
Název periodika
BIOINFORMATICS
ISSN
—
e-ISSN
1367-4811
Svazek periodika
41
Číslo periodika v rámci svazku
8
Stát vydavatele periodika
GB - Spojené království Velké Británie a Severního Irska
Počet stran výsledku
11
Strana od-do
"btaf449-1"-"btaf449-11"
Kód UT WoS článku
001554701100001
EID výsledku v databázi Scopus
2-s2.0-105013893294