Details

Vpliv medmodalne destilacije na učenje temeljnih modelov za opazovanje Zemlje
ID Tkalčič, Lana (Author), ID Čehovin Zajc, Luka (Mentor) More about this mentor... This link opens in a new window, ID Wolf, Filip (Comentor)

.pdfPDF - Presentation file, Download (2,98 MB)
MD5: BFACF726AEC2A2D1075A6CF7647B0272

Abstract
Temeljni modeli za opazovanje Zemlje se vse pogosteje učijo samonadzorovano, saj je označenih satelitskih posnetkov malo. Ogrodje DEO učenca uči z dvema učiteljema hkrati, in sicer z multispektralnim učiteljem, ki je eksponentno drseče povprečje učenca ter z zamrznjenim optičnim temeljnim modelom, pri čemer se celoten izhodni vektor učenca poravnava z optičnim učiteljem.Ker ta pozna le tri od desetih spektralnih pasov, ni jasno, koliko h kakovosti naučenih predstavitev sploh prispeva. V diplomski nalogi merimo prav ta prispevek. Kot mehanizem, s katerim optični nadzor nadzorovano zmanjšujemo, uvedemo maskiranje, ki izhodni vektor razdeli na dva dela. Prvi vstopi v destilacijo iz optičnega učitelja, drugi pa se strukturira izključno s samonadzorovano izgubo. Predučili smo šest modelov, od izhodiščnega, pri katerem se poravnava celoten vektor, do modela brez optične destilacije, in jih ovrednotili na treh klasifikacijskih ter enem segmentacijskem naboru zbirke GEO-Bench. Pokazali smo, da zmanjševanje optičnega nadzora na učni signal dejansko vpliva, saj destilacijski izgubi z večanjem maskiranja monotono padata. Vpliv je majhen, kar pripisujemo izrazni moči projekcijskih glav, ki omejitev na vhodu v veliki meri prevzamejo. Na podrejenih nalogah se prispevek optične destilacije med nabori razlikuje. Pri naboru m-eurosat uspešnost z odvzemanjem optičnega nadzora dosledno pada, pri naboru m-so2sat je najboljši model tisti s četrtino poravnanih dimenzij, pri naboru m-bigearthnet pa sta najvišje uvrščena modela z najmanj optičnega nadzora. Da bi optična destilacija kakovost naučenih predstavitev izboljšala, na naši postavitvi torej ne moremo pokazati, hkrati pa se model brez nje nikjer izrazito ne poslabša.

Language:Slovenian
Keywords:samonadzorovano učenje, destilacija znanja, opazovanje Zemlje, latentni prostor
Work type:Bachelor thesis/paper
Organization:FRI - Faculty of Computer and Information Science
Year:2026
PID:20.500.12556/RUL-187770 This link opens in a new window
Publication date in RUL:14.09.2026
Views:115
Downloads:15
Metadata:XML DC-XML DC-RDF
:
Copy citation
Share:Bookmark and Share

Secondary language

Language:English
Title:The effect of cross-modal distillation on training foundation models for Earth observation
Abstract:
Foundation models for Earth observation are increasingly trained in a self-supervised manner, since labeled satellite imagery is scarce. The DEO framework trains a student with two teachers at once, namely a multispectral teacher, which is an exponential moving average of the student, and a frozen optical foundation model, whereby the student's entire output vector is aligned with the optical teacher. Since the latter knows only three of the ten spectral bands, it is unclear how much it actually contributes to the quality of the learned representations. In this thesis we measure exactly that contribution. As a mechanism for reducing optical supervision in a controlled manner, we introduce masking that splits the output vector into two parts. The first enters distillation from the optical teacher, while the second is structured exclusively by the self-supervised loss. We pretrained six models, ranging from the baseline, in which the entire vector is aligned, to a model without optical distillation, and evaluated them on three classification and one segmentation benchmark from the GEO-Bench collection. We showed that reducing optical supervision does affect the training signal, since both distillation losses decrease monotonically with increasing masking. The effect is small, which we attribute to the expressive power of the projection heads, which largely take on the input constraint. On downstream tasks the contribution of optical distillation differs between benchmarks. On m-eurosat, performance decreases consistently as optical supervision is removed, on m-so2sat the best model is the one with a quarter of the dimensions aligned, and on m-bigearthnet the highest scores are obtained by the models with the least optical supervision. We therefore cannot show that optical distillation improves the quality of the learned representations in our setup, while at the same time the model without it does not degrade noticeably anywhere.

Keywords:self-supervised learning, knowledge distillation, Earth observation, latent space

Similar documents

Similar works from RUL:
Similar works from other Slovenian collections:

Back