Details

Up to no good : exploiting word embeddings for an automatic extraction of candidates for a lexicon of Slovene taboo language
ID Čibej, Jaka (Author)

.pdfPDF - Presentation file, Download (413,20 KB)
MD5: 6773230F4764F129D0CAA201F4F3F83D

Abstract
We present an approach to extracting candidates to be included in an open-access lexicon of Slovene taboo language by using word embeddings compiled from different Slovene corpora and a set of offensive and pejorative seed lexemes from the Thesaurus of Modern Slovene 2.0. While many studies on taboo language rely on surveys to collect data on taboo language and its use, our evaluation shows that the method with embeddings provides a good starting point for the compilation of a more comprehensive and empirically grounded taboo language lexicon. We describe the datasets used in the experiment, the process of extraction and its results, as well as the advantages and disadvantages of this method. From a set of approximately 120 Slovene seed lexemes, the initial analysis of extracted candidates resulted in 1,260 relevant lexemes. We briefly discuss potential future steps in the development of the lexicon in the context of other machine-readable Slovene language resources, such as the Digital Dictionary Database of Slovene. The extraction method is language-independent and can be directly applied to other languages.

Language:English
Keywords:taboo language, automatic extraction, embeddings, Slovene
Work type:Other
Typology:1.08 - Published Scientific Conference Contribution
Organization:FRI - Faculty of Computer and Information Science
FF - Faculty of Arts
Publication version:Version of Record
Year:2025
Number of pages:Str. 203-222
PID:20.500.12556/RUL-176621 This link opens in a new window
UDC:81'322
ISSN on article:2533-5626
COBISS.SI-ID:258417411 This link opens in a new window
Publication date in RUL:05.12.2025
Views:478
Downloads:122
Metadata:XML DC-XML DC-RDF
:
Copy citation
Share:Bookmark and Share

Record is a part of a proceedings

Title:eLex 2025
COBISS.SI-ID:258410499 This link opens in a new window

Record is a part of a journal

Title:Electronic lexicography in the 21st century. Proceedings of eLex ... conference
Shortened title:Electron. lexicogr. 21st cent., Proc. eLex ... conf.
Publisher:Lexical Computing
ISSN:2533-5626
COBISS.SI-ID:1537552579 This link opens in a new window

Licences

License:CC BY-SA 4.0, Creative Commons Attribution-ShareAlike 4.0 International
Link:http://creativecommons.org/licenses/by-sa/4.0/
Description:This Creative Commons license is very similar to the regular Attribution license, but requires the release of all derivative works under this same license.

Secondary language

Language:Slovenian
Keywords:tabujevsko besedišče, strojno luščenje, vložitve, korpusi, slovenščina

Projects

Funder:ARIS - Slovenian Research and Innovation Agency
Project number:GC-0002
Name:Veliki jezikovni modeli za digitalno humanistiko

Funder:EC - European Commission
Project number:C3.K8.IB
Name:Adaptive Natural Language Processing with Large Language Models
Acronym:PoVeJMo

Funder:ARIS - Slovenian Research and Innovation Agency
Project number:P6-0411-2019
Name:Jezikovni viri in tehnologije za slovenski jezik

Funder:ARIS - Slovenian Research and Innovation Agency
Project number:I0-E004
Name:CLARIN

Similar documents

Similar works from RUL:
Similar works from other Slovenian collections:

Back