<?xml version="1.0"?>
<metadata xmlns:xsi="http://www.w3.org/2001/XMLSchema-instance" xmlns:dc="http://purl.org/dc/elements/1.1/"><dc:title>Question-answering from old sources with large language models</dc:title><dc:creator>Tavchioski,	Ilija	(Avtor)
	</dc:creator><dc:creator>Robnik Šikonja,	Marko	(Mentor)
	</dc:creator><dc:subject>Natural Language Processing</dc:subject><dc:subject>Historical language</dc:subject><dc:subject>Machine Learning</dc:subject><dc:subject>Large Language Models</dc:subject><dc:subject>Question Answering</dc:subject><dc:description>With the expansion of artificial intelligence and the field of natural language
processing, researchers and corporations trained their models on huge
text corpora. However, a large portion of historical sources and knowledge
remain underutilized due to significant differences in vocabulary, linguistic
structure, and writing styles. In this work, we address historical texts
and documents in less-resourced languages, focused on Slovenian. By using
a digitized corpus of documents in historical Slovenian, we generated a
QA (question-answering) dataset using Slovene large language model GaMS
(Generative Model for Slovene). We supported our research on historical
Slovenian with several methodologies such as fine-tuned large language models,
PageIndex RAG (Retrieval-Augmented Generation) and a RAG approach
with a hybrid retriever, expanded with different embeddings such as Sentence
BERT, F2LLM and our own fine-tuned model. The results show that the
GaMS3 model is the most suited for generating a good QA dataset from historical
data and is the best performing model for question answering, while
the hybrid retriever enhanced with embeddings calculated from our own finetuned
model was best suited for retrieval tasks.</dc:description><dc:date>2026</dc:date><dc:date>2026-05-05 09:35:04</dc:date><dc:type>Magistrsko delo/naloga</dc:type><dc:identifier>182246</dc:identifier><dc:identifier>VisID: 38461</dc:identifier><dc:identifier>COBISS_ID: 277475587</dc:identifier><dc:language>sl</dc:language></metadata>
