<?xml version="1.0"?>
<metadata xmlns:xsi="http://www.w3.org/2001/XMLSchema-instance" xmlns:dc="http://purl.org/dc/elements/1.1/"><dc:title>Evaluation of binary classification models for the discovery of Software-as-a-Service usage</dc:title><dc:creator>DELEVA,	ROSANA	(Avtor)
	</dc:creator><dc:creator>Štruc,	Vitomir	(Mentor)
	</dc:creator><dc:subject>Software-as-a-Service</dc:subject><dc:subject>binary classification</dc:subject><dc:subject>natural language processing</dc:subject><dc:subject>text processing</dc:subject><dc:subject>transformer</dc:subject><dc:description>The exponential growth of corporate data necessitates efficient management and analysis. Software-as-a-Service (SaaS) applications have become prevalent, playing a pivotal role in modern business operations. However, accurately identifying SaaS usage within vast data streams remains a challenge. This thesis investigates the efficacy of various machine learning approaches for developing a binary classifier capable of distinguishing SaaS from non-SaaS traffic based solely on textual data strings.
The rise of big data within organizations presents both opportunities and challenges. While valuable insights can be gleaned from this data, extracting meaningful information requires robust techniques. SaaS services, with their inherent scalability and ease of use, have become the preferred choice for many business functions. 
This thesis proposes a novel approach to SaaS service identification by employing machine learning techniques. We will explore the effectiveness of several well-established algorithms, including Naive Bayes (NB) and Support Vector Machines (SVM), alongside cutting-edge transformer models like RoBERTa, XLM- RoBERTa, and the recently unveiled GPT-4. These models are trained on a meticulously curated dataset consisting of labeled text strings, representing both SaaS and non-SaaS traffic.
Through rigorous testing and evaluation, we aim to establish the optimal approach for further developing and improving our existing SaaS matching service. Our preliminary investigations suggest promising results, with certain models exhibiting a remarkable ability to differentiate between SaaS and non-SaaS data strings. 
By identifying the most effective method for developing a binary classifier, this thesis aims to contribute to the advancement of SaaS service identification and reduce the reliance on human intervention that has been necessary thus far.</dc:description><dc:date>2024</dc:date><dc:date>2024-09-20 14:25:01</dc:date><dc:type>Diplomsko delo/naloga</dc:type><dc:identifier>162293</dc:identifier><dc:identifier>VisID: 62632</dc:identifier><dc:identifier>COBISS_ID: 208717315</dc:identifier><dc:language>sl</dc:language></metadata>
