Foundation models for Earth observation are increasingly trained in a self-supervised manner, since labeled satellite imagery is scarce. The DEO framework trains a student with two teachers at once, namely a multispectral teacher, which is an exponential moving average of the student, and a frozen optical foundation model, whereby the student's entire output vector is aligned with the optical teacher. Since the latter knows only three of the ten spectral bands, it is unclear how much it actually contributes to the quality of the learned representations. In this thesis we measure exactly that contribution. As a mechanism for reducing optical supervision in a controlled manner, we introduce masking that splits the output vector into two parts. The first enters distillation from the optical teacher, while the second is structured exclusively by the self-supervised loss. We pretrained six models, ranging from the baseline, in which the entire vector is aligned, to a model without optical distillation, and evaluated them on three classification and one segmentation benchmark from the GEO-Bench collection. We showed that reducing optical supervision does affect the training signal, since both distillation losses decrease monotonically with increasing masking. The effect is small, which we attribute to the expressive power of the projection heads, which largely take on the input constraint. On downstream tasks the contribution of optical distillation differs between benchmarks. On m-eurosat, performance decreases consistently as optical supervision is removed, on m-so2sat the best model is the one with a quarter of the dimensions aligned, and on m-bigearthnet the highest scores are obtained by the models with the least optical supervision. We therefore cannot show that optical distillation improves the quality of the learned representations in our setup, while at the same time the model without it does not degrade noticeably anywhere.
|