Your browser does not allow JavaScript!
JavaScript is necessary for the proper functioning of this website. Please enable JavaScript or use a modern browser.
Repository of the University of Ljubljana
Open Science Slovenia
Open Science
DiKUL
slv
|
eng
Search
Advanced
New in RUL
About RUL
In numbers
Help
Sign in
Details
Slovene word-sense disambiguation and novelty detection with dictionary examples
ID
Škvorc, Tadej
(
Author
),
ID
Robnik Šikonja, Marko
(
Author
)
PDF - Presentation file,
Download
(1,28 MB)
MD5: 5368F2D93622A34E6519963EC2B59226
URL - Source URL, Visit
https://www.sciencedirect.com/science/article/pii/S0169023X26000856
Image galllery
Abstract
Many less-resourced languages struggle with a lack of large, task-specific datasets that are required for solving relevant tasks with modern transformer-based large language models (LLMs). On the other hand, many linguistic resources, such as dictionaries, are rarely used in this context despite their large information contents. We show how LLMs can be used to extend existing language resources in less-resourced languages for two important tasks: word-sense disambiguation (WSD) and threshold-based novelty detection (ND), a task similar to word-sense induction (WSI). We approach the two tasks through the related but much more accessible word-in-context (WiC) task where, given a pair of sentences and a target word, a classification model is tasked with predicting whether the sense of a given word differs between sentences. We demonstrate that a well-trained model for this task can distinguish between different word senses and can be adapted to solve the WSD and ND tasks. The advantage of using the WiC task, instead of directly predicting senses, is that the WiC task does not need pre-constructed sense inventories with a sufficient number of examples for each sense, which are rarely available in less-resourced languages. We show that sentence pairs for the WiC task can be successfully generated from dictionary examples using LLMs. The resulting prediction models outperform existing models on WiC, WSD, and ND tasks. We demonstrate our methodology on the Slovene language, where a monolingual dictionary is available, but word-sense resources are tiny.
Language:
English
Keywords:
large language models
,
word-sense disambiguation
,
novelty detection
,
word-in-context
,
dictionary examples
Work type:
Article
Typology:
1.01 - Original Scientific Article
Organization:
FRI - Faculty of Computer and Information Science
Publication status:
Published
Publication version:
Version of Record
Year:
2026
Number of pages:
15 str.
Numbering:
Vol. 166, art. 102638
PID:
20.500.12556/RUL-187762
UDC:
004.89:81'322
ISSN on article:
0169-023X
DOI:
10.1016/j.datak.2026.102638
COBISS.SI-ID:
290903811
Publication date in RUL:
14.09.2026
Views:
13
Downloads:
2
Metadata:
Cite this work
Plain text
BibTeX
EndNote XML
EndNote/Refer
RIS
ABNT
ACM Ref
AMA
APA
Chicago 17th Author-Date
Harvard
IEEE
ISO 690
MLA
Vancouver
:
Copy citation
Share:
Record is a part of a journal
Title:
Data & knowledge engineering
Shortened title:
Data knowl. eng.
Publisher:
Elsevier
ISSN:
0169-023X
COBISS.SI-ID:
25315584
Licences
License:
CC BY 4.0, Creative Commons Attribution 4.0 International
Link:
http://creativecommons.org/licenses/by/4.0/
Description:
This is the standard Creative Commons license that gives others maximum freedom to do what they want with the work as long as they credit the author.
Secondary language
Language:
Slovenian
Keywords:
veliki jezikovni modeli
,
razdvoumljanje pomenov besed
,
detekcija novih pomenov
,
beseda-v-kontekstu
,
slovarski primeri
Projects
Funder:
ARIS - Slovenian Research and Innovation Agency
Project number:
P6-0411-2019
Name:
Jezikovni viri in tehnologije za slovenski jezik
Funder:
ARIS - Slovenian Research and Innovation Agency
Project number:
V5-2297-2022
Name:
Medijska krajina v Sloveniji med pluralizacijo in homogenizacijo
Funder:
ARIS - Slovenian Research and Innovation Agency
Project number:
L2-50070-2023
Name:
Tehnike vektorskih vložitev za medijske aplikacije
Funder:
ARIS - Slovenian Research and Innovation Agency
Project number:
GC-0002-2024
Name:
Veliki jezikovni modeli za digitalno humanistiko
Funder:
EC - European Commission
Project number:
101186647
Name:
Centre of Excellence in Artificial Intelligence for Digital Humanities
Acronym:
AI4DH
Funder:
EC - European Commission
Project number:
C3.K8.IB
Name:
Adaptive Natural Language Processing with Large Language Models
Acronym:
PoVeJMo
Similar documents
Similar works from RUL:
Similar works from other Slovenian collections:
Back