Using structured libraries, selection, and machine learning to rapidly explore the sequence space of a fluorescent deoxyribozyme
Identifikátory výsledku
Kód výsledku v IS VaVaI
<a href="https://www.isvavai.cz/riv?ss=detail&h=RIV%2F61388963%3A_____%2F25%3A00643722" target="_blank" >RIV/61388963:_____/25:00643722 - isvavai.cz</a>
Nalezeny alternativní kódy
RIV/00216208:11310/25:10507205 RIV/60461373:22310/25:43933629
Výsledek na webu
<a href="https://doi.org/10.1093/nar/gkaf1348" target="_blank" >https://doi.org/10.1093/nar/gkaf1348</a>
DOI - Digital Object Identifier
<a href="http://dx.doi.org/10.1093/nar/gkaf1348" target="_blank" >10.1093/nar/gkaf1348</a>
Alternativní jazyky
Jazyk výsledku
angličtina
Název v původním jazyce
Using structured libraries, selection, and machine learning to rapidly explore the sequence space of a fluorescent deoxyribozyme
Popis výsledku v původním jazyce
Finding ways to more comprehensively explore the sequence space of complex functional motifs is an important and unresolved question in nucleic acid engineering. Standard approaches use libraries in which a single variant of a motif is randomly mutagenized at a low level. This provides comprehensive coverage of sequence space over short mutational distances, but only limited information about more distant variants. Here we describe a new approach that uses libraries made up of sequences consistent with the multiple constraints of a desired target motif. Functional variants are rapidly identified in a single round of selection followed by high-throughput sequencing, and rules relating sequence to function are elucidated using machine learning. This method was tested using a fluorescent deoxyribozyme recently discovered in our group called Aurora. Single-step selections showed that a secondary structure library based on Aurora contained similar to 40-fold more unique catalytic sequences than one generated by random mutagenesis. Furthermore, models developed by machine learning could quantitatively predict read numbers and identify the most active variants using small subsets of sequences as training sets. By combining secondary structure libraries, selection, and machine learning in this way, sequence space can be explored far more quickly and efficiently than in standard approaches.
Název v anglickém jazyce
Using structured libraries, selection, and machine learning to rapidly explore the sequence space of a fluorescent deoxyribozyme
Popis výsledku anglicky
Finding ways to more comprehensively explore the sequence space of complex functional motifs is an important and unresolved question in nucleic acid engineering. Standard approaches use libraries in which a single variant of a motif is randomly mutagenized at a low level. This provides comprehensive coverage of sequence space over short mutational distances, but only limited information about more distant variants. Here we describe a new approach that uses libraries made up of sequences consistent with the multiple constraints of a desired target motif. Functional variants are rapidly identified in a single round of selection followed by high-throughput sequencing, and rules relating sequence to function are elucidated using machine learning. This method was tested using a fluorescent deoxyribozyme recently discovered in our group called Aurora. Single-step selections showed that a secondary structure library based on Aurora contained similar to 40-fold more unique catalytic sequences than one generated by random mutagenesis. Furthermore, models developed by machine learning could quantitatively predict read numbers and identify the most active variants using small subsets of sequences as training sets. By combining secondary structure libraries, selection, and machine learning in this way, sequence space can be explored far more quickly and efficiently than in standard approaches.
Klasifikace
Druh
J<sub>imp</sub> - Článek v periodiku v databázi Web of Science
CEP obor
—
OECD FORD obor
10608 - Biochemistry and molecular biology
Návaznosti výsledku
Projekt
Výsledek vznikl pri realizaci vícero projektů. Více informací v záložce Projekty.
Návaznosti
I - Institucionalni podpora na dlouhodoby koncepcni rozvoj vyzkumne organizace
Ostatní
Rok uplatnění
2025
Kód důvěrnosti údajů
S - Úplné a pravdivé údaje o projektu nepodléhají ochraně podle zvláštních právních předpisů
Údaje specifické pro druh výsledku
Název periodika
Nucleic Acids Research
ISSN
0305-1048
e-ISSN
1362-4962
Svazek periodika
53
Číslo periodika v rámci svazku
22
Stát vydavatele periodika
US - Spojené státy americké
Počet stran výsledku
11
Strana od-do
gkaf1348
Kód UT WoS článku
001637985400001
EID výsledku v databázi Scopus
2-s2.0-105024587762