Segmentation of satellite imagery represents a crucial stage in remote sensing
data analysis for urban planning, environmental monitoring, and disaster
response. Despite the success of discriminative models, such approaches often
fail to capture spatial coherence, boundary uncertainty, and image parts
ozadjethat could be construed as multiple classes. In this thesis, we address
these limitations by introducing a generative framework that reformulates
satellite image segmentation as the synthesis of a segmentation mask in latent
space via rectified flow. The generative process is conditioned on features
from the pretrained visual encoders DINOv3 and DEO. We evaluate the
proposed method on three publicly available datasets: SpaceNetv1, GeoBench
Chesapeake, and GeoBench Cashew, covering binary building detection as
well as multiclass segmentation of urban and agricultural land cover from
satellite imagery. The proposed method achieves a mean IoU of 70.40, ranking
second among the compared approaches. On the multispectral GeoBench
Cashew dataset, it achieves the best result among all compared methods,
with a macro-IoU of 69.67. We demonstrate that encoder adaptation and the
choice of the source distribution affect the quality of the results.
|