Using structured libraries, selection, and machine learning to rapidly explore the sequence space of a fluorescent deoxyribozyme
The result's identifiers
Result code in IS VaVaI
<a href="https://www.isvavai.cz/riv?ss=detail&h=RIV%2F61388963%3A_____%2F25%3A00643722" target="_blank" >RIV/61388963:_____/25:00643722 - isvavai.cz</a>
Alternative codes found
RIV/00216208:11310/25:10507205 RIV/60461373:22310/25:43933629
Result on the web
<a href="https://doi.org/10.1093/nar/gkaf1348" target="_blank" >https://doi.org/10.1093/nar/gkaf1348</a>
DOI - Digital Object Identifier
<a href="http://dx.doi.org/10.1093/nar/gkaf1348" target="_blank" >10.1093/nar/gkaf1348</a>
Alternative languages
Result language
angličtina
Original language name
Using structured libraries, selection, and machine learning to rapidly explore the sequence space of a fluorescent deoxyribozyme
Original language description
Finding ways to more comprehensively explore the sequence space of complex functional motifs is an important and unresolved question in nucleic acid engineering. Standard approaches use libraries in which a single variant of a motif is randomly mutagenized at a low level. This provides comprehensive coverage of sequence space over short mutational distances, but only limited information about more distant variants. Here we describe a new approach that uses libraries made up of sequences consistent with the multiple constraints of a desired target motif. Functional variants are rapidly identified in a single round of selection followed by high-throughput sequencing, and rules relating sequence to function are elucidated using machine learning. This method was tested using a fluorescent deoxyribozyme recently discovered in our group called Aurora. Single-step selections showed that a secondary structure library based on Aurora contained similar to 40-fold more unique catalytic sequences than one generated by random mutagenesis. Furthermore, models developed by machine learning could quantitatively predict read numbers and identify the most active variants using small subsets of sequences as training sets. By combining secondary structure libraries, selection, and machine learning in this way, sequence space can be explored far more quickly and efficiently than in standard approaches.
Czech name
—
Czech description
—
Classification
Type
J<sub>imp</sub> - Article in a specialist periodical, which is included in the Web of Science database
CEP classification
—
OECD FORD branch
10608 - Biochemistry and molecular biology
Result continuities
Project
Result was created during the realization of more than one project. More information in the Projects tab.
Continuities
I - Institucionalni podpora na dlouhodoby koncepcni rozvoj vyzkumne organizace
Others
Publication year
2025
Confidentiality
S - Úplné a pravdivé údaje o projektu nepodléhají ochraně podle zvláštních právních předpisů
Data specific for result type
Name of the periodical
Nucleic Acids Research
ISSN
0305-1048
e-ISSN
1362-4962
Volume of the periodical
53
Issue of the periodical within the volume
22
Country of publishing house
US - UNITED STATES
Number of pages
11
Pages from-to
gkaf1348
UT code for WoS article
001637985400001
EID of the result in the Scopus database
2-s2.0-105024587762