Synthetic Aperture Radar (SAR) enables Earth surface observation regardless of illumination conditions or cloud cover. Due to the process of image
acquisition and the characteristics of radar backscatter, SAR imagery differs substantially from optical imagery, raising the question of the extent to
which domain-specific adaptations are necessary when applying deep learning methods. In this thesis, we examine this question in the context of bitemporal food detection using Sentinel-1 imagery from the OMBRIA dataset.
We adapt the BTC architecture to SAR imagery and integrate it with the
Copernicus-FM foundation model. We experimentally evaluate the impact of
different data augmentation techniques and examine whether their physical
plausibility with respect to SAR image acquisition is reflected in model performance. The results show no clear relationship between the physical plausibility of the augmentations and their effectiveness. The best-performing
combination, which also includes transformations that are physically questionable in the context of SAR, achieves an average F1 = 0.8529. An additional comparison of encoder configurations shows that the best performance
is achieved by a Swin-T encoder pretrained on optical imagery
|