TmProt 1.0: A Web Server for Predicting Protein Melting Temperatures from Sequences*
Identifikátory výsledku
Kód výsledku v IS VaVaI
<a href="https://www.isvavai.cz/riv?ss=detail&h=RIV%2F00216224%3A14310%2F25%3A00143722" target="_blank" >RIV/00216224:14310/25:00143722 - isvavai.cz</a>
Výsledek na webu
—
DOI - Digital Object Identifier
—
Alternativní jazyky
Jazyk výsledku
angličtina
Název v původním jazyce
TmProt 1.0: A Web Server for Predicting Protein Melting Temperatures from Sequences*
Popis výsledku v původním jazyce
Accurate prediction of protein melting temperature (Tm) is essential for understanding protein stability and guiding bioengineering efforts. We developed a transformer-based model tailored for Tm prediction, leveraging the large-scale Meltome Atlas dataset for initial training. Given that the Meltome Atlas provides only proxy Tm values due to its experimental design, we further refined our model using the BRENDA database, which reports experimentally validated Tm values. This two-stage approach enabled the transfer of knowledge from a broad but noisy dataset to a smaller, higher-quality dataset. We investigated two strategies for this transfer: (i) training a custom transformer model from scratch on low-quality data followed by fine-tuning with high-quality data, and (ii) adapting a pre-trained foundation model (ESM-2) using parameter-efficient fine-tuning via LoRA. Additionally, we compared these methods with an MLP model trained on ESM-2 embeddings. (popis)
Název v anglickém jazyce
TmProt 1.0: A Web Server for Predicting Protein Melting Temperatures from Sequences*
Popis výsledku anglicky
Accurate prediction of protein melting temperature (Tm) is essential for understanding protein stability and guiding bioengineering efforts. We developed a transformer-based model tailored for Tm prediction, leveraging the large-scale Meltome Atlas dataset for initial training. Given that the Meltome Atlas provides only proxy Tm values due to its experimental design, we further refined our model using the BRENDA database, which reports experimentally validated Tm values. This two-stage approach enabled the transfer of knowledge from a broad but noisy dataset to a smaller, higher-quality dataset. We investigated two strategies for this transfer: (i) training a custom transformer model from scratch on low-quality data followed by fine-tuning with high-quality data, and (ii) adapting a pre-trained foundation model (ESM-2) using parameter-efficient fine-tuning via LoRA. Additionally, we compared these methods with an MLP model trained on ESM-2 embeddings. (popis)
Klasifikace
Druh
R - Software
CEP obor
—
OECD FORD obor
10201 - Computer sciences, information science, bioinformathics (hardware development to be 2.2, social aspect to be 5.8)
Návaznosti výsledku
Projekt
<a href="/cs/project/LM2023049" target="_blank" >LM2023049: Český národní uzel Evropské sítě infrastruktur klinického výzkumu</a><br>
Návaznosti
P - Projekt vyzkumu a vyvoje financovany z verejnych zdroju (s odkazem do CEP)
Ostatní
Rok uplatnění
2025
Kód důvěrnosti údajů
U - Předmět řešení projektu je utajovanou skutečností podle zvláštních právních předpisů nebo je skutečností, jejíž zveřejnění by mohlo ohrozit činnost zpravodajské služby. Údaje o projektu jsou upraveny tak, aby byly zveřejnitelné
Údaje specifické pro druh výsledku
Interní identifikační kód produktu
01_2025_LL
Technické parametry
Doposud je využíváno interně k výzkumným účelům.
Ekonomické parametry
V budoucnu je plánována nabídka potencionálním partnerům z komerční sféry.
IČO vlastníka výsledku
00216224
Název vlastníka
Masarykova univerzita