Your browser does not allow JavaScript!
JavaScript is necessary for the proper functioning of this website. Please enable JavaScript or use a modern browser.
Repository of the University of Ljubljana
Open Science Slovenia
Open Science
DiKUL
slv
|
eng
Search
Advanced
New in RUL
About RUL
In numbers
Help
Sign in
Details
Up to no good : exploiting word embeddings for an automatic extraction of candidates for a lexicon of Slovene taboo language
ID
Čibej, Jaka
(
Author
)
PDF - Presentation file,
Download
(413,20 KB)
MD5: 6773230F4764F129D0CAA201F4F3F83D
Image galllery
Abstract
We present an approach to extracting candidates to be included in an open-access lexicon of Slovene taboo language by using word embeddings compiled from different Slovene corpora and a set of offensive and pejorative seed lexemes from the Thesaurus of Modern Slovene 2.0. While many studies on taboo language rely on surveys to collect data on taboo language and its use, our evaluation shows that the method with embeddings provides a good starting point for the compilation of a more comprehensive and empirically grounded taboo language lexicon. We describe the datasets used in the experiment, the process of extraction and its results, as well as the advantages and disadvantages of this method. From a set of approximately 120 Slovene seed lexemes, the initial analysis of extracted candidates resulted in 1,260 relevant lexemes. We briefly discuss potential future steps in the development of the lexicon in the context of other machine-readable Slovene language resources, such as the Digital Dictionary Database of Slovene. The extraction method is language-independent and can be directly applied to other languages.
Language:
English
Keywords:
taboo language
,
automatic extraction
,
embeddings
,
Slovene
Work type:
Other
Typology:
1.08 - Published Scientific Conference Contribution
Organization:
FRI - Faculty of Computer and Information Science
FF - Faculty of Arts
Publication version:
Version of Record
Year:
2025
Number of pages:
Str. 203-222
PID:
20.500.12556/RUL-176621
UDC:
81'322
ISSN on article:
2533-5626
COBISS.SI-ID:
258417411
Publication date in RUL:
05.12.2025
Views:
478
Downloads:
122
Metadata:
Cite this work
Plain text
BibTeX
EndNote XML
EndNote/Refer
RIS
ABNT
ACM Ref
AMA
APA
Chicago 17th Author-Date
Harvard
IEEE
ISO 690
MLA
Vancouver
:
Copy citation
Share:
Record is a part of a proceedings
Title:
eLex 2025
COBISS.SI-ID:
258410499
Record is a part of a journal
Title:
Electronic lexicography in the 21st century. Proceedings of eLex ... conference
Shortened title:
Electron. lexicogr. 21st cent., Proc. eLex ... conf.
Publisher:
Lexical Computing
ISSN:
2533-5626
COBISS.SI-ID:
1537552579
Licences
License:
CC BY-SA 4.0, Creative Commons Attribution-ShareAlike 4.0 International
Link:
http://creativecommons.org/licenses/by-sa/4.0/
Description:
This Creative Commons license is very similar to the regular Attribution license, but requires the release of all derivative works under this same license.
Secondary language
Language:
Slovenian
Keywords:
tabujevsko besedišče
,
strojno luščenje
,
vložitve
,
korpusi
,
slovenščina
Projects
Funder:
ARIS - Slovenian Research and Innovation Agency
Project number:
GC-0002
Name:
Veliki jezikovni modeli za digitalno humanistiko
Funder:
EC - European Commission
Project number:
C3.K8.IB
Name:
Adaptive Natural Language Processing with Large Language Models
Acronym:
PoVeJMo
Funder:
ARIS - Slovenian Research and Innovation Agency
Project number:
P6-0411-2019
Name:
Jezikovni viri in tehnologije za slovenski jezik
Funder:
ARIS - Slovenian Research and Innovation Agency
Project number:
I0-E004
Name:
CLARIN
Similar documents
Similar works from RUL:
Similar works from other Slovenian collections:
Back