Unsupervised Semantic Segmentation of Urban Scenes viaCross-Modal Distillation
Identifikátory výsledku
Kód výsledku v IS VaVaI
<a href="https://www.isvavai.cz/riv?ss=detail&h=RIV%2F68407700%3A21230%2F25%3A00388541" target="_blank" >RIV/68407700:21230/25:00388541 - isvavai.cz</a>
Nalezeny alternativní kódy
RIV/68407700:21730/25:00388541
Výsledek na webu
<a href="https://doi.org/10.1007/s11263-024-02320-3" target="_blank" >https://doi.org/10.1007/s11263-024-02320-3</a>
DOI - Digital Object Identifier
<a href="http://dx.doi.org/10.1007/s11263-024-02320-3" target="_blank" >10.1007/s11263-024-02320-3</a>
Alternativní jazyky
Jazyk výsledku
angličtina
Název v původním jazyce
Unsupervised Semantic Segmentation of Urban Scenes viaCross-Modal Distillation
Popis výsledku v původním jazyce
Semantic image segmentation models typically require extensive pixel-wise annotations, which are costly to obtain and proneto biases. Our work investigates learning semantic segmentation in urban scenes without any manual annotation. We proposea novel method for learning pixel-wise semantic segmentation using raw, uncurated data from vehicle-mounted camerasand LiDAR sensors, thus eliminating the need for manual labeling. Our contributions are as follows. First, we develop anovel approach for cross-modal unsupervised learning of semantic segmentation by leveraging synchronized LiDAR andimage data. A crucial element of our method is the integration of an object proposal module that examines the LiDARpoint cloud to generate proposals for spatially consistent objects. Second, we demonstrate that these 3D object proposalscan be aligned with corresponding images and effectively grouped into semantically meaningful pseudo-classes. Third, weintroduce a cross-modal distillation technique that utilizes image data partially annotated with the learnt pseudo-classes totrain a transformer-based model for semantic image segmentation. Fourth, we demonstrate further significant improvementsof our approach by extending the proposed model using a teacher-student distillation with an exponential moving average andincorporating soft targets from the teacher. We show the generalization capabilities of our method by testing on four differenttesting datasets (Cityscapes, Dark Zurich, Nighttime Driving, and ACDC) without any fine-tuning. We present an in-depthexperimental analysis of the proposed model including results when using another pre-training dataset, per-class and pixelaccuracy results, confusion matrices, PCA visualization, k-NN evaluation, ablations of the number of clusters and LiDAR’sdensity, supervised finetuning as well as additional qualitative results and their analysis.
Název v anglickém jazyce
Unsupervised Semantic Segmentation of Urban Scenes viaCross-Modal Distillation
Popis výsledku anglicky
Semantic image segmentation models typically require extensive pixel-wise annotations, which are costly to obtain and proneto biases. Our work investigates learning semantic segmentation in urban scenes without any manual annotation. We proposea novel method for learning pixel-wise semantic segmentation using raw, uncurated data from vehicle-mounted camerasand LiDAR sensors, thus eliminating the need for manual labeling. Our contributions are as follows. First, we develop anovel approach for cross-modal unsupervised learning of semantic segmentation by leveraging synchronized LiDAR andimage data. A crucial element of our method is the integration of an object proposal module that examines the LiDARpoint cloud to generate proposals for spatially consistent objects. Second, we demonstrate that these 3D object proposalscan be aligned with corresponding images and effectively grouped into semantically meaningful pseudo-classes. Third, weintroduce a cross-modal distillation technique that utilizes image data partially annotated with the learnt pseudo-classes totrain a transformer-based model for semantic image segmentation. Fourth, we demonstrate further significant improvementsof our approach by extending the proposed model using a teacher-student distillation with an exponential moving average andincorporating soft targets from the teacher. We show the generalization capabilities of our method by testing on four differenttesting datasets (Cityscapes, Dark Zurich, Nighttime Driving, and ACDC) without any fine-tuning. We present an in-depthexperimental analysis of the proposed model including results when using another pre-training dataset, per-class and pixelaccuracy results, confusion matrices, PCA visualization, k-NN evaluation, ablations of the number of clusters and LiDAR’sdensity, supervised finetuning as well as additional qualitative results and their analysis.
Klasifikace
Druh
J<sub>imp</sub> - Článek v periodiku v databázi Web of Science
CEP obor
—
OECD FORD obor
10201 - Computer sciences, information science, bioinformathics (hardware development to be 2.2, social aspect to be 5.8)
Návaznosti výsledku
Projekt
<a href="/cs/project/EF15_003%2F0000468" target="_blank" >EF15_003/0000468: Inteligentní strojové vnímání</a><br>
Návaznosti
P - Projekt vyzkumu a vyvoje financovany z verejnych zdroju (s odkazem do CEP)<br>S - Specificky vyzkum na vysokych skolach
Ostatní
Rok uplatnění
2025
Kód důvěrnosti údajů
S - Úplné a pravdivé údaje o projektu nepodléhají ochraně podle zvláštních právních předpisů
Údaje specifické pro druh výsledku
Název periodika
International Journal of Computer Vision
ISSN
0920-5691
e-ISSN
1573-1405
Svazek periodika
133
Číslo periodika v rámci svazku
6
Stát vydavatele periodika
NL - Nizozemsko
Počet stran výsledku
23
Strana od-do
3519-3541
Kód UT WoS článku
001396170300001
EID výsledku v databázi Scopus
2-s2.0-85217156465