<?xml version="1.0"?>
<metadata xmlns:xsi="http://www.w3.org/2001/XMLSchema-instance" xmlns:dc="http://purl.org/dc/elements/1.1/"><dc:title>A semi-automatic video object segmentation method</dc:title><dc:creator>Pelhan,	Jer	(Avtor)
	</dc:creator><dc:creator>Kristan,	Matej	(Mentor)
	</dc:creator><dc:subject>convolutional neural network</dc:subject><dc:subject>video object segmentation</dc:subject><dc:subject>video object tracking</dc:subject><dc:description>Visual object tracking has recently shifted towards target segmentation, which has increased the demand for video datasets with objects segmented in each frame. However, manually obtaining large segmented video datasets is time-consuming and costly. We address this problem by introducing a Semi-supervised Annotation by Tracking algorithm (SAT), which is specialized for target segmentation specifically for visual object tracking domain with minimal user input. The annotation pipeline is split into two modules. The anchor frame segmentation module predicts a segmentation mask by few (approximately four) user clicks on the object of interest. The module is used to segment the target in a subset of frames, anchors, throughout the sequence. Then a mask propagation module propagates the segmentation masks from the anchors to the in-between frames. On the VOT dataset, SAT achieves an IoU of 73% already at 5% of user annotated frames and outperforms the winner of the DAVIS2020 challenge IVOS and the winner of DAVIS2018 challenge IVS by 40% and 67%, respectively and shortens the annotation time by 98%. On the DAVIS interactive challenges, SAT performs comparably to the state-of-the-art in video object segmentation.</dc:description><dc:date>2021</dc:date><dc:date>2021-05-13 11:45:01</dc:date><dc:type>Diplomsko delo/naloga</dc:type><dc:identifier>127022</dc:identifier><dc:identifier>VisID: 31741</dc:identifier><dc:identifier>COBISS_ID: 63110659</dc:identifier><dc:language>sl</dc:language></metadata>
