<?xml version="1.0"?>
<metadata xmlns:xsi="http://www.w3.org/2001/XMLSchema-instance" xmlns:dc="http://purl.org/dc/elements/1.1/"><dc:title>Semi-supervised relation extraction corpus construction and models creation for under-resourced languages</dc:title><dc:creator>Knez,	Timotej	(Avtor)
	</dc:creator><dc:creator>Štravs,	Miha	(Avtor)
	</dc:creator><dc:creator>Žitnik,	Slavko	(Avtor)
	</dc:creator><dc:subject>relation extraction</dc:subject><dc:subject>semi-supervised learning</dc:subject><dc:subject>Slovene language</dc:subject><dc:description>The goal of relation extraction is to recognize head and tail entities in a document and determine a relation between them. While a lot of progress was made in solving automated relation extraction in widely used languages such as English, the use of these methods for under-resourced languages and domains is limited due to the lack of training data. In this work, we present a pipeline using distant supervision for constructing a relation extraction corpus in an arbitrary language. The corpus construction combines Wikipedia documents in the target language with relations in the WikiData knowledge graph. We demonstrate the process by constructing a new corpus for relation extraction in the Slovene language. Our corpus captures 20 unique relation types. The final corpus contains 811,032 relations annotated in 244,437 sentences. We use the corpus to train models using three architectures and evaluate them on the task of Slovene relation extraction. We achieve comparable performance to approaches on English data.</dc:description><dc:date>2025</dc:date><dc:date>2025-08-27 14:39:00</dc:date><dc:type>Članek v reviji</dc:type><dc:identifier>171514</dc:identifier><dc:identifier>UDK: 004.65:81'322</dc:identifier><dc:identifier>ISSN pri članku: 2078-2489</dc:identifier><dc:identifier>DOI: 10.3390/info16020143</dc:identifier><dc:identifier>COBISS_ID: 226450691</dc:identifier><dc:language>sl</dc:language></metadata>
