Accurate predictions of enzymatic biochemistry as an enabler for generation of de-novo sequences
Identifikátory výsledku
Kód výsledku v IS VaVaI
<a href="https://www.isvavai.cz/riv?ss=detail&h=RIV%2F68407700%3A21730%2F24%3A00380848" target="_blank" >RIV/68407700:21730/24:00380848 - isvavai.cz</a>
Výsledek na webu
<a href="https://openreview.net/pdf?id=tWL5iQIlMQ" target="_blank" >https://openreview.net/pdf?id=tWL5iQIlMQ</a>
DOI - Digital Object Identifier
—
Alternativní jazyky
Jazyk výsledku
angličtina
Název v původním jazyce
Accurate predictions of enzymatic biochemistry as an enabler for generation of de-novo sequences
Popis výsledku v původním jazyce
Terpene synthases (TPSs) generate the scaffolds of the largest class of natural products, including several first-line medicines. The amount of available TPS protein sequences is increasing exponentially, but computational characterization of their function remains an unsolved challenge. We assembled a curated dataset of one thousand characterized TPS reactions and developed a method to devise highly accurate machine-learning models for functional annotation in a low-data regime. Our models significantly outperform existing methods for TPS detection and substrate prediction. By applying the models to large protein sequence databases, we discovered seven TPS enzymes previously undetected by state-of-the-art computational tools and experimentally confirmed their activity. Furthermore, we discovered a new TPS structural domain and distinct subtypes of previously known domains. Our work demonstrates the potential of machine learning to speed up the discovery and characterization of novel TPSs. Furthermore, in-silico functional annotations provide the ML community with a large dataset of pseudo-labeled exemplary TPS sequences. The accurate models for TPS detection and substrate prediction can serve as oracles to check the presence of desired biochemical activity in the generated sequences. We envision the published dataset of exemplary TPS sequences and the accurate TPS-annotation models to boost the generation of de-novo enzymatic TPS sequences.
Název v anglickém jazyce
Accurate predictions of enzymatic biochemistry as an enabler for generation of de-novo sequences
Popis výsledku anglicky
Terpene synthases (TPSs) generate the scaffolds of the largest class of natural products, including several first-line medicines. The amount of available TPS protein sequences is increasing exponentially, but computational characterization of their function remains an unsolved challenge. We assembled a curated dataset of one thousand characterized TPS reactions and developed a method to devise highly accurate machine-learning models for functional annotation in a low-data regime. Our models significantly outperform existing methods for TPS detection and substrate prediction. By applying the models to large protein sequence databases, we discovered seven TPS enzymes previously undetected by state-of-the-art computational tools and experimentally confirmed their activity. Furthermore, we discovered a new TPS structural domain and distinct subtypes of previously known domains. Our work demonstrates the potential of machine learning to speed up the discovery and characterization of novel TPSs. Furthermore, in-silico functional annotations provide the ML community with a large dataset of pseudo-labeled exemplary TPS sequences. The accurate models for TPS detection and substrate prediction can serve as oracles to check the presence of desired biochemical activity in the generated sequences. We envision the published dataset of exemplary TPS sequences and the accurate TPS-annotation models to boost the generation of de-novo enzymatic TPS sequences.
Klasifikace
Druh
D - Stať ve sborníku
CEP obor
—
OECD FORD obor
10201 - Computer sciences, information science, bioinformathics (hardware development to be 2.2, social aspect to be 5.8)
Návaznosti výsledku
Projekt
—
Návaznosti
I - Institucionalni podpora na dlouhodoby koncepcni rozvoj vyzkumne organizace
Ostatní
Rok uplatnění
2024
Kód důvěrnosti údajů
S - Úplné a pravdivé údaje o projektu nepodléhají ochraně podle zvláštních právních předpisů
Údaje specifické pro druh výsledku
Název statě ve sborníku
Proceeding The Twelfth International Conference on Learning Representations (ICLR 2024)
ISBN
9781713898658
ISSN
—
e-ISSN
—
Počet stran výsledku
3
Strana od-do
—
Název nakladatele
International Conference on Learning Representations
Místo vydání
—
Místo konání akce
Vídeň
Datum konání akce
7. 5. 2024
Typ akce podle státní příslušnosti
WRD - Celosvětová akce
Kód UT WoS článku
—