Unsupervised Semantic Segmentation of Urban Scenes viaCross-Modal Distillation
The result's identifiers
Result code in IS VaVaI
<a href="https://www.isvavai.cz/riv?ss=detail&h=RIV%2F68407700%3A21230%2F25%3A00388541" target="_blank" >RIV/68407700:21230/25:00388541 - isvavai.cz</a>
Alternative codes found
RIV/68407700:21730/25:00388541
Result on the web
<a href="https://doi.org/10.1007/s11263-024-02320-3" target="_blank" >https://doi.org/10.1007/s11263-024-02320-3</a>
DOI - Digital Object Identifier
<a href="http://dx.doi.org/10.1007/s11263-024-02320-3" target="_blank" >10.1007/s11263-024-02320-3</a>
Alternative languages
Result language
angličtina
Original language name
Unsupervised Semantic Segmentation of Urban Scenes viaCross-Modal Distillation
Original language description
Semantic image segmentation models typically require extensive pixel-wise annotations, which are costly to obtain and proneto biases. Our work investigates learning semantic segmentation in urban scenes without any manual annotation. We proposea novel method for learning pixel-wise semantic segmentation using raw, uncurated data from vehicle-mounted camerasand LiDAR sensors, thus eliminating the need for manual labeling. Our contributions are as follows. First, we develop anovel approach for cross-modal unsupervised learning of semantic segmentation by leveraging synchronized LiDAR andimage data. A crucial element of our method is the integration of an object proposal module that examines the LiDARpoint cloud to generate proposals for spatially consistent objects. Second, we demonstrate that these 3D object proposalscan be aligned with corresponding images and effectively grouped into semantically meaningful pseudo-classes. Third, weintroduce a cross-modal distillation technique that utilizes image data partially annotated with the learnt pseudo-classes totrain a transformer-based model for semantic image segmentation. Fourth, we demonstrate further significant improvementsof our approach by extending the proposed model using a teacher-student distillation with an exponential moving average andincorporating soft targets from the teacher. We show the generalization capabilities of our method by testing on four differenttesting datasets (Cityscapes, Dark Zurich, Nighttime Driving, and ACDC) without any fine-tuning. We present an in-depthexperimental analysis of the proposed model including results when using another pre-training dataset, per-class and pixelaccuracy results, confusion matrices, PCA visualization, k-NN evaluation, ablations of the number of clusters and LiDAR’sdensity, supervised finetuning as well as additional qualitative results and their analysis.
Czech name
—
Czech description
—
Classification
Type
J<sub>imp</sub> - Article in a specialist periodical, which is included in the Web of Science database
CEP classification
—
OECD FORD branch
10201 - Computer sciences, information science, bioinformathics (hardware development to be 2.2, social aspect to be 5.8)
Result continuities
Project
<a href="/en/project/EF15_003%2F0000468" target="_blank" >EF15_003/0000468: Intelligent Machine Perception</a><br>
Continuities
P - Projekt vyzkumu a vyvoje financovany z verejnych zdroju (s odkazem do CEP)<br>S - Specificky vyzkum na vysokych skolach
Others
Publication year
2025
Confidentiality
S - Úplné a pravdivé údaje o projektu nepodléhají ochraně podle zvláštních právních předpisů
Data specific for result type
Name of the periodical
International Journal of Computer Vision
ISSN
0920-5691
e-ISSN
1573-1405
Volume of the periodical
133
Issue of the periodical within the volume
6
Country of publishing house
NL - THE KINGDOM OF THE NETHERLANDS
Number of pages
23
Pages from-to
3519-3541
UT code for WoS article
001396170300001
EID of the result in the Scopus database
2-s2.0-85217156465