Leveraging large language models for literature-driven prioritization of protein binding pockets
The result's identifiers
Result code in IS VaVaI
<a href="https://www.isvavai.cz/riv?ss=detail&h=RIV%2F61989592%3A15310%2F25%3A73635235" target="_blank" >RIV/61989592:15310/25:73635235 - isvavai.cz</a>
Result on the web
<a href="https://academic.oup.com/bioinformatics/article/41/8/btaf449/8225722" target="_blank" >https://academic.oup.com/bioinformatics/article/41/8/btaf449/8225722</a>
DOI - Digital Object Identifier
<a href="http://dx.doi.org/10.1093/bioinformatics/btaf449" target="_blank" >10.1093/bioinformatics/btaf449</a>
Alternative languages
Result language
angličtina
Original language name
Leveraging large language models for literature-driven prioritization of protein binding pockets
Original language description
Motivation Accurately identifying and prioritizing protein binding pockets is a foundational element of small-molecule drug discovery. Defining these known pockets currently relies on a laborious manual process of extracting key residue data from selected publications, reconciling inconsistent terminology, and independently computing volumetric representations. This manual curation to ensure biological relevance is time-consuming, error-prone, and represents a major bottleneck for efficient, high-throughput drug discovery. Results We present a novel approach for the identification and prioritization of protein binding pockets for small molecules by combining geometric pocket detection with large language models (LLMs). Our method leverages Fpocket to generate candidate pockets, which are then validated against published experimental data extracted from research articles using LLM with a series of prompts fine-tuned to identify and extract residue-level information associated with experimentally confirmed binding sites. We developed a curated benchmark dataset of diverse proteins and associated literature to train and evaluate the LLM's performance in paper relevance assessment and pocket extraction. Availability and implementation The developed benchmark dataset and methodology are freely available at the GitHub repository (https://github.com/receptor-ai/LLM-benchmark-dataset) and Zenodo (DOI: 10.5281/zenodo.15798647).
Czech name
—
Czech description
—
Classification
Type
J<sub>imp</sub> - Article in a specialist periodical, which is included in the Web of Science database
CEP classification
—
OECD FORD branch
10403 - Physical chemistry
Result continuities
Project
—
Continuities
I - Institucionalni podpora na dlouhodoby koncepcni rozvoj vyzkumne organizace
Others
Publication year
2025
Confidentiality
S - Úplné a pravdivé údaje o projektu nepodléhají ochraně podle zvláštních právních předpisů
Data specific for result type
Name of the periodical
BIOINFORMATICS
ISSN
—
e-ISSN
1367-4811
Volume of the periodical
41
Issue of the periodical within the volume
8
Country of publishing house
GB - UNITED KINGDOM
Number of pages
11
Pages from-to
"btaf449-1"-"btaf449-11"
UT code for WoS article
001554701100001
EID of the result in the Scopus database
2-s2.0-105013893294