<?xml version="1.0"?>
<rdf:RDF xmlns:rdf="http://www.w3.org/1999/02/22-rdf-syntax-ns#" xmlns:dc="http://purl.org/dc/elements/1.1/"><rdf:Description rdf:about="https://repozitorij.uni-lj.si/IzpisGradiva.php?id=175070"><dc:title>A method for object detection by text prompting</dc:title><dc:creator>Rot,	Žiga	(Avtor)
	</dc:creator><dc:creator>Kristan,	Matej	(Mentor)
	</dc:creator><dc:creator>Pelhan,	Jer	(Komentor)
	</dc:creator><dc:subject>object detection</dc:subject><dc:subject>computer vision</dc:subject><dc:subject>deep learning</dc:subject><dc:description>The increasing diversity and scale of object detection datasets have highlighted the limitations of closed-set detectors with fixed vocabularies. Open-set object detection addresses this by enabling the detection of arbitrary classes via text prompts. Grounding DINO is a prominent zero-shot detector, but its training code is not fully open-source, and its implementation is outdated. In this work, we reimplement Grounding DINO, achieving ~20% speedup, and extend it for text-based object counting, though these modifications do not consistently improve FSCD-147 performance. To enable training from scratch, we optimize the model further, achieving an additional ~30% speedup, and develop an edge-oriented variant inspired by a closed-source model. Training on 1.3 million images from scratch, we evaluate the model on COCO and LVIS datasets, where it performs comparably to other open-source models, but remains below the closed-source baseline, likely due to a much smaller training set.</dc:description><dc:date>2025</dc:date><dc:date>2025-10-14 13:55:01</dc:date><dc:type>Magistrsko delo/naloga</dc:type><dc:identifier>175070</dc:identifier><dc:language>sl</dc:language></rdf:Description></rdf:RDF>
