This thesis addresses scene-level image retrieval from partial
freehand sketches, where a small number of strokes may admit several
possible interpretations of the scene. To model this ambiguity, we
propose a conditional generative model based on rectified flow that
generates multiple final representations for the same partial sketch
from different random initial states in a shared sketch-image
embedding space. We use the geometric dispersion of these
representations and disagreement between their retrieval results as
indirect indicators of model uncertainty. Experiments on FS-COCO show
that disagreement between generated representations generally
decreases as the sketch becomes more complete, while individual
queries may retain relatively high dispersion even when the sketch is
nearly or fully completed. Qualitative analysis shows that geometric
dispersion may, but does not necessarily, result in disagreement
between retrieval outputs. We additionally evaluate the generated
representations for conventional image retrieval and compare them
with direct use of the original embeddings and deterministic mapping
approaches. These approaches can achieve stronger retrieval
performance, but return only a single final representation for each
sketch and therefore do not support disagreement analysis through
sampling. The results indicate that generative modelling with
multiple samples can provide an additional signal about model
behaviour for ambiguous queries beyond the retrieval result itself.
|