Details

Segmentacija objektov skozi transformacije z uporabo modela SAM
ID Jezovšek, Gašper (Author), ID Lukežič, Alan (Mentor) More about this mentor... This link opens in a new window

.pdfPDF - Presentation file, Download (8,28 MB)
MD5: DF0EFA23E99DDA39140CC310BE8A9F36

Abstract
Segmentacija in sledenje objektov v videu sta zreli področji, sodobni modeli, kot je SAM2, pa objektom sledijo predvsem na podlagi videza. Kadar objekt med posnetkom bistveno spremeni obliko, na primer med rezanjem, lomljenjem ali gnetenjem, razpade na več delov in sledenje pogosto odpove. V diplomskem delu prilagodimo modela družine Segment Anything za segmentacijo objektov skozi tovrstne transformacije. Model SAM2 doučimo na podatkovni zbirki VOST prek obstoječega učnega cevovoda, za model SAM3 pa razvijemo lasten postopek videoučenja sledilnika, pri katerem gradienti tečejo skozi časovni pomnilnik, tako da se ta med učenjem prilagodi nalogi. Doučenje izboljša oba modela: povprečno prekrivanje (J) naraste z 0,471 na 0,539 pri SAM2 in z 0,536 na 0,614 pri SAM3, pri čemer je izboljšava največja prav v tistih delih posnetka, kjer se objekt preoblikuje. Doučeni model dosega tudi najboljši rezultat na tekmovanju VOTSt2025. Ovrednotimo tudi prenosljivost na običajne objekte in pokažemo, da specializacija na razpadajoče objekte zniža natančnost pri njih.

Language:Slovenian
Keywords:segmentacija objektov, sledenje objektom v videu, transformacije objektov, SAM, doučenje, VOST
Work type:Bachelor thesis/paper
Typology:2.11 - Undergraduate Thesis
Organization:FRI - Faculty of Computer and Information Science
Year:2026
PID:20.500.12556/RUL-185790 This link opens in a new window
COBISS.SI-ID:289079811 This link opens in a new window
Publication date in RUL:20.08.2026
Views:168
Downloads:93
Metadata:XML DC-XML DC-RDF
:
Copy citation
Share:Bookmark and Share

Secondary language

Language:English
Title:Object segmentation through transformations using the SAM model
Abstract:
Video object segmentation and tracking are mature fields, yet modern models such as SAM2 track objects mainly by their appearance. When an object substantially changes shape during a sequence, for example while being cut, broken, or kneaded, it falls apart into several pieces and tracking often fails. In this thesis we adapt models of the Segment Anything family for object segmentation through such transformations. We fine-tune SAM2 on the VOST dataset using the model’s existing training pipeline, while for SAM3 we develop our own video-training procedure for the tracker, in which gradients flow through the temporal memory so that it adapts to the task during training. Fine-tuning improves both models: the mean region overlap (J) rises from 0.471 to 0.539 for SAM2 and from 0.536 to 0.614 for SAM3, with the largest gains precisely on the sequence sections where transformations happen. Our fine-tuned model also achieves the best result on the VOTSt2025 challenge. We also evaluate transfer to ordinary objects and show that specialising for disintegrating objects reduces accuracy on them.

Keywords:object segmentation, video object tracking, object transformations, SAM, fine-tuning, VOST

Similar documents

Similar works from RUL:
Similar works from other Slovenian collections:

Back