<?xml version="1.0"?>
<rdf:RDF xmlns:rdf="http://www.w3.org/1999/02/22-rdf-syntax-ns#" xmlns:dc="http://purl.org/dc/elements/1.1/"><rdf:Description rdf:about="https://repozitorij.uni-lj.si/IzpisGradiva.php?id=169302"><dc:title>Augmentation of positive and unlabeled data using generative adversarial networks</dc:title><dc:creator>Papič,	Aleš	(Avtor)
	</dc:creator><dc:creator>Bosnić,	Zoran	(Mentor)
	</dc:creator><dc:creator>Kononenko,	Igor	(Komentor)
	</dc:creator><dc:subject>machine learning</dc:subject><dc:subject>deep learning</dc:subject><dc:subject>generative adversarial networks</dc:subject><dc:subject>data augmentation</dc:subject><dc:subject>positive and unlabeled learning</dc:subject><dc:subject>binary classification</dc:subject><dc:description>In an era characterized by rapid technological advancements, generative artificial intelligence is steadily finding its way into consumer electronics. As content generation becomes more effortless, the consequent fast data growth poses significant challenges for data processing, often reliant on human labor.

This thesis explores positive and unlabeled learning as a strategy to reduce the cost of data labeling. The primary advantage of positive and unlabeled learning is its effectiveness when negative data are either unavailable or too diverse to label directly. By leveraging both positive and unlabeled data, positive and unlabeled learning utilizes all available information, offering greater robustness and generalization compared to methods that rely solely on positive data. We propose a novel Conditional Generative Positive and Unlabeled (CGenPU) framework, which trains a binary classifier to differentiate between known positive and unknown negative examples. 

To effectively train the classifier, we need to simultaneously train a generator to generate both positive and negative training examples for the classifier. Since existing loss functions require labeled examples from all relevant classes, we developed our own loss function to address this limitation. Specifically, to enable the classifier to effectively discriminate between positive and negative examples, we introduce a novel auxiliary loss that facilitates learning from positive and unlabeled datasets. The soundness of our approach is demonstrated through theoretical analysis. We apply CGenPU to binary image classification tasks using multiple benchmark datasets, such as MNIST and CIFAR-10. Evaluation shows superior performance in digit recognition and object classification tasks. However, the weak nature of auxiliary loss indicates stability and overfitting issues.

To address these limitations, we propose Positively Dense Example Weighting (PosiDEW), which calculates weights for training examples, improving class balance in sampled batch data. Additionally, we extend the auxiliary loss with a regularization term, which prevents overfitting by slowing down the classifier's learning. Evaluation demonstrates that the proposed improvements enable CGenPU to more effectively learn the distribution of positive and negative data. Furthermore, these refinements do not negatively impact training time while significantly improve classification accuracy.

Additionally, we propose a novel polyp detection pipeline to address the slow and tedious labeling process. CGenPU trains a classifier to generate polyp segmentation masks, which are then postprocessed to determine polyp locations on an image. Evaluation reveals comparable performance with existing approaches, although it has not yet reach the level of state-of-the-art techniques.</dc:description><dc:date>2025</dc:date><dc:date>2025-05-22 15:35:02</dc:date><dc:type>Doktorsko delo/naloga</dc:type><dc:identifier>169302</dc:identifier><dc:language>sl</dc:language></rdf:Description></rdf:RDF>
