<?xml version="1.0"?>
<rdf:RDF xmlns:rdf="http://www.w3.org/1999/02/22-rdf-syntax-ns#" xmlns:dc="http://purl.org/dc/elements/1.1/"><rdf:Description rdf:about="https://repozitorij.uni-lj.si/IzpisGradiva.php?id=141868"><dc:title>Cross-lingual word embeddings for knowledge transfer in less-represented languages</dc:title><dc:creator>Škvorc,	Tadej	(Avtor)
	</dc:creator><dc:creator>Robnik Šikonja,	Marko	(Mentor)
	</dc:creator><dc:subject>natural language processing</dc:subject><dc:subject>deep learning</dc:subject><dc:subject>neural networks</dc:subject><dc:subject>machine learning</dc:subject><dc:subject>text embeddings</dc:subject><dc:subject>knowledge transfer</dc:subject><dc:subject>idiom detection</dc:subject><dc:subject>conference scheduling</dc:subject><dc:subject>multilingual embeddings</dc:subject><dc:subject>contextual embeddings</dc:subject><dc:description>Neural networks and deep learning have led to big advances in the field of natural language processing. However, many such techniques rely on large, manually annotated datasets, which are not always available, particularly for less popular tasks and less-resourced languages. In our thesis, we show how text embeddings and knowledge transfer can be used to improve upon existing state-of-the-art approaches and make them viable for less-resourced languages. We demonstrate the developed methodological novelties on two advanced use cases.

In the first part of the thesis, we focus on idiom detection. We use contextual and multilingual text embeddings and develop a novel method that outperforms existing approaches. Our method is capable of detecting idioms that do not appear in the training set, which is a major advantage over existing methods. We evaluate our approach on a novel Slovene dataset and a multilingual dataset of 20 languages. We show that our approach is capable of generalizing between closely related languages (i.e. Slovene and Croatian), that it is able to function with only a small amount of training data, and that we can use it on a related task of metaphor detection with the use of  knowledge transfer.

In the second part of the thesis, we present a method for automatic conference scheduling. The method arranges paper presentations into a schedule of a scientific conference, minimizing  overlaps between presentations with similar topics. We use text and graph-based features to find similar papers and arrange them into a schedule using a novel algorithm based on constrained clustering and optimization. We evaluate our approach both on synthetic data and multiple real-world conferences in English and Slovene. Our approach does not require a labelled dataset and makes use of multilingual embeddings, making it suitable for less-represented languages.

Our work shows how text embeddings and knowledge transfer can be used to improve current NLP approaches on less-resourced languages, sometimes removing the need for large annotated datasets that has traditionally been a problem for deep learning approaches.</dc:description><dc:date>2022</dc:date><dc:date>2022-10-10 10:15:00</dc:date><dc:type>Doktorsko delo/naloga</dc:type><dc:identifier>141868</dc:identifier><dc:language>sl</dc:language></rdf:Description></rdf:RDF>
