Improving machine learning-based bitewing segmentation with synthetic data
Identifikátory výsledku
Kód výsledku v IS VaVaI
<a href="https://www.isvavai.cz/riv?ss=detail&h=RIV%2F00064165%3A_____%2F25%3A10497991" target="_blank" >RIV/00064165:_____/25:10497991 - isvavai.cz</a>
Nalezeny alternativní kódy
RIV/00216208:11110/25:10497991
Výsledek na webu
<a href="https://verso.is.cuni.cz/pub/verso.fpl?fname=obd_publikace_handle&handle=IV8jIDQdVO" target="_blank" >https://verso.is.cuni.cz/pub/verso.fpl?fname=obd_publikace_handle&handle=IV8jIDQdVO</a>
DOI - Digital Object Identifier
<a href="http://dx.doi.org/10.1016/j.jdent.2025.105679" target="_blank" >10.1016/j.jdent.2025.105679</a>
Alternativní jazyky
Jazyk výsledku
angličtina
Název v původním jazyce
Improving machine learning-based bitewing segmentation with synthetic data
Popis výsledku v původním jazyce
Objectives: Class imbalance in datasets is one of the challenges of machine learning (ML) in medical image analysis. We employed synthetic data to overcome class imbalance when segmenting bitewing radiographs as an exemplary task for using ML. Methods: After segmenting bitewings into classes, i.e. dental structures, restorations, and background, the pixellevel representation of implants in the training set (1543 bitewings) and testing set (177 bitewings) was 0.03 % and 0.07 %, respectively. A diffusion model and a generative adversarial network (pix2pix) were used to generate a dataset synthetically enriched in implants. A U-Net segmentation model was trained on (1) the original dataset, (2) the synthetic dataset, (3) on the synthetic dataset and fine-tuned on the original dataset, or (4) on a dataset which was na & iuml;vely oversampled with images containing implants. Results: U-Net trained on the original dataset was unable to segment implants in the testing set. Model performance was significantly improved by na & iuml;ve over-sampling, achieving the highest precision. The model trained only on synthetic data performed worse than na & iuml;ve over-sampling in all metrics, but with fine-tuning on original data, it resulted in the highest Dice score, recall, F1 score and ROC AUC, respectively. The performance on other classes than implants was similar for all strategies except training only on synthetic data, which tended to perform worse. Conclusions: The use of synthetic data alone may deteriorate the performance of segmentation models. However, fine-tuning on original data could significantly enhance model performance, especially for heavily underrepresented classes. Clinical significance: This study explored the use of synthetic data to enhance segmentation of bitewing radiographs, focusing on underrepresented classes like implants. Pre-training on synthetic data followed by finetuning on original data yielded the best results, highlighting the potential of synthetic data to advance AIdriven dental imaging and ultimately support clinical decision-making.
Název v anglickém jazyce
Improving machine learning-based bitewing segmentation with synthetic data
Popis výsledku anglicky
Objectives: Class imbalance in datasets is one of the challenges of machine learning (ML) in medical image analysis. We employed synthetic data to overcome class imbalance when segmenting bitewing radiographs as an exemplary task for using ML. Methods: After segmenting bitewings into classes, i.e. dental structures, restorations, and background, the pixellevel representation of implants in the training set (1543 bitewings) and testing set (177 bitewings) was 0.03 % and 0.07 %, respectively. A diffusion model and a generative adversarial network (pix2pix) were used to generate a dataset synthetically enriched in implants. A U-Net segmentation model was trained on (1) the original dataset, (2) the synthetic dataset, (3) on the synthetic dataset and fine-tuned on the original dataset, or (4) on a dataset which was na & iuml;vely oversampled with images containing implants. Results: U-Net trained on the original dataset was unable to segment implants in the testing set. Model performance was significantly improved by na & iuml;ve over-sampling, achieving the highest precision. The model trained only on synthetic data performed worse than na & iuml;ve over-sampling in all metrics, but with fine-tuning on original data, it resulted in the highest Dice score, recall, F1 score and ROC AUC, respectively. The performance on other classes than implants was similar for all strategies except training only on synthetic data, which tended to perform worse. Conclusions: The use of synthetic data alone may deteriorate the performance of segmentation models. However, fine-tuning on original data could significantly enhance model performance, especially for heavily underrepresented classes. Clinical significance: This study explored the use of synthetic data to enhance segmentation of bitewing radiographs, focusing on underrepresented classes like implants. Pre-training on synthetic data followed by finetuning on original data yielded the best results, highlighting the potential of synthetic data to advance AIdriven dental imaging and ultimately support clinical decision-making.
Klasifikace
Druh
J<sub>imp</sub> - Článek v periodiku v databázi Web of Science
CEP obor
—
OECD FORD obor
30208 - Dentistry, oral surgery and medicine
Návaznosti výsledku
Projekt
—
Návaznosti
V - Vyzkumna aktivita podporovana z jinych verejnych zdroju
Ostatní
Rok uplatnění
2025
Kód důvěrnosti údajů
S - Úplné a pravdivé údaje o projektu nepodléhají ochraně podle zvláštních právních předpisů
Údaje specifické pro druh výsledku
Název periodika
Journal of Dentistry
ISSN
0300-5712
e-ISSN
1879-176X
Svazek periodika
156
Číslo periodika v rámci svazku
May
Stát vydavatele periodika
GB - Spojené království Velké Británie a Severního Irska
Počet stran výsledku
7
Strana od-do
105679
Kód UT WoS článku
001447050400001
EID výsledku v databázi Scopus
2-s2.0-86000488979