Details

Slovene word-sense disambiguation and novelty detection with dictionary examples
ID Škvorc, Tadej (Author), ID Robnik Šikonja, Marko (Author)

.pdfPDF - Presentation file, Download (1,28 MB)
MD5: 5368F2D93622A34E6519963EC2B59226
URLURL - Source URL, Visit https://www.sciencedirect.com/science/article/pii/S0169023X26000856 This link opens in a new window

Abstract
Many less-resourced languages struggle with a lack of large, task-specific datasets that are required for solving relevant tasks with modern transformer-based large language models (LLMs). On the other hand, many linguistic resources, such as dictionaries, are rarely used in this context despite their large information contents. We show how LLMs can be used to extend existing language resources in less-resourced languages for two important tasks: word-sense disambiguation (WSD) and threshold-based novelty detection (ND), a task similar to word-sense induction (WSI). We approach the two tasks through the related but much more accessible word-in-context (WiC) task where, given a pair of sentences and a target word, a classification model is tasked with predicting whether the sense of a given word differs between sentences. We demonstrate that a well-trained model for this task can distinguish between different word senses and can be adapted to solve the WSD and ND tasks. The advantage of using the WiC task, instead of directly predicting senses, is that the WiC task does not need pre-constructed sense inventories with a sufficient number of examples for each sense, which are rarely available in less-resourced languages. We show that sentence pairs for the WiC task can be successfully generated from dictionary examples using LLMs. The resulting prediction models outperform existing models on WiC, WSD, and ND tasks. We demonstrate our methodology on the Slovene language, where a monolingual dictionary is available, but word-sense resources are tiny.

Language:English
Keywords:large language models, word-sense disambiguation, novelty detection, word-in-context, dictionary examples
Work type:Article
Typology:1.01 - Original Scientific Article
Organization:FRI - Faculty of Computer and Information Science
Publication status:Published
Publication version:Version of Record
Year:2026
Number of pages:15 str.
Numbering:Vol. 166, art. 102638
PID:20.500.12556/RUL-187762 This link opens in a new window
UDC:004.89:81'322
ISSN on article:0169-023X
DOI:10.1016/j.datak.2026.102638 This link opens in a new window
COBISS.SI-ID:290903811 This link opens in a new window
Publication date in RUL:14.09.2026
Views:13
Downloads:2
Metadata:XML DC-XML DC-RDF
:
Copy citation
Share:Bookmark and Share

Record is a part of a journal

Title:Data & knowledge engineering
Shortened title:Data knowl. eng.
Publisher:Elsevier
ISSN:0169-023X
COBISS.SI-ID:25315584 This link opens in a new window

Licences

License:CC BY 4.0, Creative Commons Attribution 4.0 International
Link:http://creativecommons.org/licenses/by/4.0/
Description:This is the standard Creative Commons license that gives others maximum freedom to do what they want with the work as long as they credit the author.

Secondary language

Language:Slovenian
Keywords:veliki jezikovni modeli, razdvoumljanje pomenov besed, detekcija novih pomenov, beseda-v-kontekstu, slovarski primeri

Projects

Funder:ARIS - Slovenian Research and Innovation Agency
Project number:P6-0411-2019
Name:Jezikovni viri in tehnologije za slovenski jezik

Funder:ARIS - Slovenian Research and Innovation Agency
Project number:V5-2297-2022
Name:Medijska krajina v Sloveniji med pluralizacijo in homogenizacijo

Funder:ARIS - Slovenian Research and Innovation Agency
Project number:L2-50070-2023
Name:Tehnike vektorskih vložitev za medijske aplikacije

Funder:ARIS - Slovenian Research and Innovation Agency
Project number:GC-0002-2024
Name:Veliki jezikovni modeli za digitalno humanistiko

Funder:EC - European Commission
Project number:101186647
Name:Centre of Excellence in Artificial Intelligence for Digital Humanities
Acronym:AI4DH

Funder:EC - European Commission
Project number:C3.K8.IB
Name:Adaptive Natural Language Processing with Large Language Models
Acronym:PoVeJMo

Similar documents

Similar works from RUL:
Similar works from other Slovenian collections:

Back