The thesis addresses the classification of images in which a part of the pixels is missing and the positions of the missing pixels are described by a binary mask, known both at training and at prediction time. As the main contribution we propose a two-phase method that merges a partial-convolution U-Net inpainter and a ResNet classifier into a single network: it first learns to inpaint, then the focus of training shifts to classification while the inpainting part keeps adapting to the classification signal, so that the fill-ins become useful for classification. On a subset of one hundred ImageNet classes at resolution 224 × 224 we establish a unified experimental environment with three mask families (square regions, Perlin noise, random pixels) and a wide range of missing-pixel ratios. Within it we compare a baseline ResNet with zero filling, a classifier with partial convolutions, the spatial graph network SGCN, the MisConv method, and an inpaint-then-classify pipeline. The proposed method achieves the highest accuracy for all mask families and all missing-pixel ratios and outperforms the strongest published method, MisConv, also on the datasets of the original paper. The results further quantify the benefit of knowing the mask at a realistic resolution and how the accuracy depends on the mask family and the missing-pixel ratio.
|