<?xml version="1.0"?>
<metadata xmlns:xsi="http://www.w3.org/2001/XMLSchema-instance" xmlns:dc="http://purl.org/dc/elements/1.1/"><dc:title>First steps toward the compilation of a safety dataset for Slovene large language models</dc:title><dc:creator>Čibej,	Jaka	(Avtor)
	</dc:creator><dc:subject>large language models</dc:subject><dc:subject>responsible artificial intelligence</dc:subject><dc:subject>safety datasets</dc:subject><dc:subject>Slovene</dc:subject><dc:description>In the paper, we present the initial preparatory phase of the compilation of a Slovene safety dataset containing harmful or offensive prompts and safe responses to them. The dataset will be used to fine-tune Slovene large language models in order to prevent unwanted model behavior and misuse by malicious actors for a diverse range of harmful activities, such as scams, toxic or offensive content generation, automated political campaigning, vandalism, and terrorism. We provide an overview of existing safety datasets for other languages and describe the different methods used to compile them, as well as the harm areas typically covered in similar datasets. We continue by listing the most frequent vulnerabilities of existing LLMs and how to take them in to account when designing a safety dataset that covers not only the general harm areas, but also those specific to Slovenia. Wep ropose a framework for the manual generation of Slovene prompts and responses based on an initial taxonomy of relevant topics, along with additional instructions to provide for more linguistic diversity with in the dataset and account forpotential frequent jailbreaks.</dc:description><dc:date>2024</dc:date><dc:date>2024-10-18 11:22:02</dc:date><dc:type>Drugo</dc:type><dc:identifier>164271</dc:identifier><dc:identifier>UDK: 81'322:004.8</dc:identifier><dc:identifier>COBISS_ID: 212026627</dc:identifier><dc:identifier>OceCobissID: 211315971</dc:identifier><dc:language>sl</dc:language></metadata>
