<?xml version="1.0"?>
<rdf:RDF xmlns:rdf="http://www.w3.org/1999/02/22-rdf-syntax-ns#" xmlns:dc="http://purl.org/dc/elements/1.1/"><rdf:Description rdf:about="https://repozitorij.uni-lj.si/IzpisGradiva.php?id=177747"><dc:title>Leveraging a morphological lexicon for a semi-automatic approach to correcting lemmas and morphosyntactic tags</dc:title><dc:creator>Čibej,	Jaka	(Avtor)
	</dc:creator><dc:creator>Munda,	Tina	(Avtor)
	</dc:creator><dc:subject>lemmatization</dc:subject><dc:subject>morphosyntactic tagging</dc:subject><dc:subject>training corpora</dc:subject><dc:subject>morphological lexicon</dc:subject><dc:subject>corpus annotation</dc:subject><dc:description>In the paper, we present a new semi-automatic approach to correcting lemmas and morpho-syntactic tags. Unlike previous manual annotation approaches for Slovene corpora, the new method contains an additional step in which tokens and their automatically assigned lemmas and  morphosyntactic  tags  are  cross-referenced  with  the  set  of  forms  included  in  the  Sloleks  Morphological Lexicon of Slovene. Based on the comparison, each token is classified into one of several annotation scenarios. The new approach has noticeably reduced the time and resources invested into annotation by eliminating a large number of redundant tasks. The advantages of this method include the possibility of dividing annotation tasks into groups consisting of simi-lar annotation problems (e.g. disambiguation of grammatical homographs). With adequate data preparation,  it  also  drastically  reduces  the  necessity  for  annotators  to  be  familiar  with  the  extensive  Multext-East  morphosyntactic  tag  set  for  Slovene,  a  restriction  that  created  a  bottleneck in the annotation process in similar annotation campaigns. The method was tested during the annotation process for the ROG Training Corpus of Spoken Slovene. In addition, we also test the scenario classification algorithm on the SUK Training Corpus of Written Slovene, which was annotated using the traditional sentence-by-sentence, token-by-token approach. We present the results and argue that the method should be used in future annotation campaigns to save resources and improve overall annotation consistency, while also discussing some of the caveats and disadvantages of the proposed approach.</dc:description><dc:date>2025</dc:date><dc:date>2026-01-06 10:22:02</dc:date><dc:type>Članek v reviji</dc:type><dc:identifier>177747</dc:identifier><dc:language>sl</dc:language></rdf:Description></rdf:RDF>
