<?xml version="1.0"?>
<rdf:RDF xmlns:rdf="http://www.w3.org/1999/02/22-rdf-syntax-ns#" xmlns:dc="http://purl.org/dc/elements/1.1/"><rdf:Description rdf:about="https://repozitorij.uni-lj.si/IzpisGradiva.php?id=176621"><dc:title>Up to no good</dc:title><dc:creator>Čibej,	Jaka	(Avtor)
	</dc:creator><dc:subject>taboo language</dc:subject><dc:subject>automatic extraction</dc:subject><dc:subject>embeddings</dc:subject><dc:subject>Slovene</dc:subject><dc:description>We present an approach to extracting candidates to be included in an open-access lexicon of Slovene taboo language by using word embeddings compiled from different Slovene corpora and a set of offensive and pejorative seed lexemes from the Thesaurus of Modern Slovene 2.0. While many studies on taboo language rely on surveys to collect data on taboo language and its use, our evaluation shows that the method with embeddings provides a good starting point for the compilation of a more comprehensive and empirically grounded taboo language lexicon. We describe the datasets used in the experiment, the process of extraction and its results, as well as the advantages and disadvantages of this method. From a set of approximately 120 Slovene seed lexemes, the initial analysis of extracted candidates resulted in 1,260 relevant lexemes. We briefly discuss potential future steps in the development of the lexicon in the context of other machine-readable Slovene language resources, such as the Digital Dictionary Database of Slovene. The extraction method is language-independent and can be directly applied to other languages.</dc:description><dc:date>2025</dc:date><dc:date>2025-12-05 12:36:39</dc:date><dc:type>Drugo</dc:type><dc:identifier>176621</dc:identifier><dc:language>sl</dc:language></rdf:Description></rdf:RDF>
