<?xml version="1.0"?>
<metadata xmlns:xsi="http://www.w3.org/2001/XMLSchema-instance" xmlns:dc="http://purl.org/dc/elements/1.1/"><dc:title>Commonsense-enhanced deep neural networks for less-resourced languages</dc:title><dc:creator>Žagar,	Aleš	(Avtor)
	</dc:creator><dc:creator>Robnik Šikonja,	Marko	(Mentor)
	</dc:creator><dc:creator>Kosem,	Iztok	(Komentor)
	</dc:creator><dc:subject>natural language processing</dc:subject><dc:subject>text summarization</dc:subject><dc:subject>text simplification</dc:subject><dc:subject>commonsense reasoning</dc:subject><dc:subject>language resources for Slovene</dc:subject><dc:subject>large language models</dc:subject><dc:description>This doctoral dissertation addresses the gap in the development of modern language technologies for Slovene, a less-resourced language that lags behind high-resource languages, particularly in the fields of text generation and complex semantic reasoning. The work combines the construction of essential data resources, the development of new methodologies for text summarization and simplification, and an in-depth analysis of the commonsense reasoning capabilities of large language models (LLMs).
In the first part, we present infrastructural improvements that serve as a foundation for training and evaluating neural models. We developed an updated version of the Corpus of Academic Slovene (KAS 2.0), which, through cleaner extraction and structural segmentation, enables the creation of dedicated datasets for long-document summarization and machine translation. For standardized evaluation of natural language understanding, we established the Slovene SuperGLUE benchmark, enabling a direct comparison of the performance of monolingual and multilingual models.
In the field of text generation, we introduce innovative approaches to address the scarcity of training data. We present a cross-lingual transfer method for summarization that leverages knowledge transfer from English models to effectively summarize Slovene news articles. We further advance the field with the development of the SloMetaSum, which, rather than seeking a universal model, employs a neural network to dynamically select the most suitable summarizer based on the content and structure of the input document. To enhance information accessibility, we designed SENTA, the first comprehensive solution for Slovene sentence simplification, combining complexity classification and generative rewriting using a fine-tuned T5 model.
The final part of the thesis focuses on commonsense reasoning. By constructing the extended KE-WSC dataset, enriched with structured ontological annotations and human explanations, we analyze the impact of explicit knowledge on model performance. The results indicate that the inclusion of explanations improves accuracy primarily in models specialized for reasoning, while also revealing a significant gap between a model's ability to solve the task and its capacity to generate high-quality justifications.
Through freely available datasets, models, and evaluation studies, this dissertation significantly contributes to the development of semantically deep and user-centric language technologies with focus on Slovene.</dc:description><dc:date>2026</dc:date><dc:date>2026-09-27 12:30:03</dc:date><dc:type>Doktorsko delo/naloga</dc:type><dc:identifier>188783</dc:identifier><dc:identifier>VisID: 37216</dc:identifier><dc:language>sl</dc:language></metadata>
