<?xml version="1.0"?>
<rdf:RDF xmlns:rdf="http://www.w3.org/1999/02/22-rdf-syntax-ns#" xmlns:dc="http://purl.org/dc/elements/1.1/"><rdf:Description rdf:about="https://repozitorij.uni-lj.si/IzpisGradiva.php?id=127413"><dc:title>Discriminative appearance models for efficient correlation-based visual object tracking</dc:title><dc:creator>Lukežič,	Alan	(Avtor)
	</dc:creator><dc:creator>Kristan,	Matej	(Mentor)
	</dc:creator><dc:subject>Visual object tracking</dc:subject><dc:subject>short-term tracking</dc:subject><dc:subject>long-term tracking</dc:subject><dc:subject>discriminative correlation filters</dc:subject><dc:subject>deformable objects</dc:subject><dc:description>Visual object tracking addresses target trajectory estimation in a video sequence given a single training example in the first frame. Diverse factors such as occlusion, illumination change, fast object or camera motion, object deformation, clutter and target disappearance make visual tracking particularly challenging. In this thesis we focus on methodological framework of discriminative correlation filters (DCFs), which shows a great potential in tracking. We propose four contributions to DCF-based tracking. The first three contributions address short-term tracking of deformable and non-compact targets, which are poorly approximated by axis-aligned bounding boxes. The last contribution addresses long-term tracking in which the target disappears and remains absent for long periods before re-appearing. The first contribution explores the problem of deformable target tracking. We propose a part-based visual model that considers the target appearance at two levels of details. At coarse level, a holistic target representation is maintained by a segmentation model combined with a DCF, while a geometrically-constrained constellation of DCFs is used for detailed representation. We formulate the per-part visual similarity terms and the inter-part geometric deformation constraints within a single spring-system-based model and propose an efficient optimization to find the maximum a posteriori solution. A drawback of the part-based models is the limited amount of deformations that the model can describe. Moreover, when the target does not deform, estimation of a large number of deformation parameters from an uncertain visual data may deteriorate tracking performance. In our second contribution, we thus explore a holistic model which applies a spatial attention mechanism to identify the target pixels during training and applies channel attention to select the features most suitable for target tracking. We propose a channel and spatial reliability discriminative correlation filter (CSRDCF). An approximate spatial attention map is generated as a color-based segmentation mask and used to constrain the support of the trained DCF. We propose an efficient optimization for the mask-constrained filter learning. Channel attention, on the other hand is estimated by inspecting the per-channel localization quality during learning. The resulting tracker runs in real-time on a CPU and attains a high degree of robustness. While the target mask estimated by traditional color-based methods may be sufficient for attention mechanism in constrained DCF learning, it is not accurate enough for representing the target location. In recent years, however, deep convolutional neural networks have been shown to generate highly accurate segmentations. In the third contribution we thus revise discriminative tracking in the context of a deep neural network. We propose a single-stage segmentation tracker (D3S), whose primary output is the target segmentation mask. The network combines a deep variant of a DCF and a nonparametric appearance model to discriminatively specialize to the selected target and produce a high-fidelity segmentation mask. The network is trained on segmentation task only, generalizes to a range of targets and achieves a state-of-the-art tracking performance. In the fourth contribution we propose a new DCF-based long-term tracker. The tracker is composed of a short-term component, responsible for frame-to-frame localization, and of a detector, responsible for image-wide target re-localization after target loss. Both the short-term component and the detector are formulated as constrained DCFs and a mechanism for efficient interaction between the two models is proposed. In addition, we propose a long-term tracking performance evaluation methodology and a benchmark. The benchmark consists of a long-term tracking dataset focusing mostly on target disappearances, a taxonomy which positions trackers on short/long-term spectrum and novel long-term tracking performance measures. The methodology and the dataset have been used as a part of the largest visual object tracking challenge VOT.</dc:description><dc:date>2021</dc:date><dc:date>2021-06-04 12:43:01</dc:date><dc:type>Doktorsko delo/naloga</dc:type><dc:identifier>127413</dc:identifier><dc:language>sl</dc:language></rdf:Description></rdf:RDF>
