Details

Universal Dependencies za slovenščino : nove smernice, ročno označeni podatki in razčlenjevalni model
ID Dobrovoljc, Kaja (Author), ID Terčon, Luka (Author), ID Ljubešić, Nikola (Author)

.pdfPDF - Presentation file, Download (517,92 KB)
MD5: 44E1790F2FA1FDA6AAF46BE71BCBE520
URLURL - Source URL, Visit https://journals.uni-lj.si/slovenscina2/article/view/12031/13792 This link opens in a new window

Abstract
Universal Dependencies (UD) je mednarodno usklajena označevalna shema za medjezikovno primerljivo oblikoslovno in skladenjsko označevanje besedil po načelih odvisnostne slovnice, ki je bila ob več kot 130 drugih svetovnih jezikih uspešno uporabljena tudi za označevanje besedil v slovenščini. V prispevku predstavimo rezultate nedavnih aktivnosti v povezavi s shemo UD znotraj projekta Razvoj slovenščine v digitalnem okolju, v okviru katerega smo obstoječo infrastrukturo nadgradili s prenovo in podrobno dokumentacijo označevalnih smernic UD za slovenščino, razširitvijo drevesnice SSJ-UD za pisno slovenščino z novimi povedmi iz korpusov ssj500k in ELEXIS-WSD, izdelavo testne množice iz besedil korpusa SentiCoref za spletni portal SloBENCH ter polavtomatsko pretvorbo oblikoslovnih oznak referenčnih učnih korpusov SUK in Janes-Tag. Na razširjeni drevesnici SSJ-UD je bil naučen tudi novi napovedni model za skladenjsko razčlenjevanje v orodju CLASSLA-Stanza, ki ga v prispevku v podporo nadaljnjim jezikoslovnim aplikacijam podrobneje ovrednotimo z vidika splošne natančnosti razčlenjevanja in najpogostejših tipov napak.

Language:Slovenian
Keywords:slovnično označeni korpusi, odvisnostna slovnica, drevesnica, skladenjsko razčlenjevanje, računalniško jezikoslovje, obdelava naravnega jezika (računalništvo)
Typology:1.01 - Original Scientific Article
Organization:FF - Faculty of Arts
Publication version:Version of Record
Publication date:01.01.2023
Year:2023
Number of pages:Str. 218-246
Numbering:Letn. 11, št. 1
PID:20.500.12556/RUL-166876 This link opens in a new window
UDC:004.9:81
ISSN on article:2335-2736
DOI:10.4312/slo2.0.2023.1.218-246 This link opens in a new window
COBISS.SI-ID:173830147 This link opens in a new window
Publication date in RUL:29.01.2025
Views:588
Downloads:221
Metadata:XML DC-XML DC-RDF
:
Copy citation
Share:Bookmark and Share

Record is a part of a journal

Title:Slovenščina 2.0 : empirične, aplikativne in interdisciplinarne raziskave
Publisher:Trojina, zavod za uporabno slovenistiko, Trojina, zavod za uporabno slovenistiko, Trojina, zavod za uporabno slovenistiko, Znanstvena založba Filozofske fakultete, Znanstvena založba Filozofske fakultete, Založba Univerze v Ljubljani
ISSN:2335-2736
COBISS.SI-ID:264547328 This link opens in a new window

Licences

License:CC BY-SA 4.0, Creative Commons Attribution-ShareAlike 4.0 International
Link:http://creativecommons.org/licenses/by-sa/4.0/
Description:This Creative Commons license is very similar to the regular Attribution license, but requires the release of all derivative works under this same license.

Secondary language

Language:English
Title:Universal dependencies for Slovenian
Abstract:
Universal Dependencies (UD) is an internationally coordinated annotation scheme for cross-linguistically comparable morphosyntactic annotation of corpora, which has been applied to more than 130 other languages world-wide, including Slovenian. In this paper, we present the results of recent ac-tivities related to Slovenian UD annotation within the Development of Slovene in a Digital Environment project. During the project, we upgraded the existing infrastructure with reviewed and detailed documentation of the Slovenian UD annotation guidelines and produced four new datasets, manually annotated in accordance with the scheme. Specifically, we expanded the SSJ-UD treebank for written Slovenian with new sentences from the ssj500k and ELEXIS-WSD corpora, and created a new hidden UD treebank based on the SentiCoref cor-pus to be used on the SloBENCH evaluation platform. In addition, the SUK and Janes-tag reference training corpora, originally annotated using the language-specific JOS annotation scheme, have been semi-automatically converted to UD part-of-speech categories and morphological features. The new version of the reference SSJ-UD treebank with more than 5,000 new sentences and double the original number of tokens was used to train a new dependency parsing model in the CLASSLA-Stanza annotation tool. This paper gives an in-depth evaluation of its performance with respect to the overall parsing per-formance, the relation-specific parsing performance and the most common types of errors produced.

Keywords:linguistic annotation, dependency grammar, treebanks, dependency parsing, natural language processing, computational linguistics, natural language processing (computer science)

Projects

Funder:Other - Other funder or multiple funders
Funding programme:Ministrstvo za kulturo Republike Slovenije
Project number:-
Name:Razvoj slovenščine v digitalnem okolju
Acronym:RSDO

Funder:ARIS - Slovenian Research and Innovation Agency
Project number:P6-0411
Name:Jezikovni viri in tehnologije za slovenski jezik

Similar documents

Similar works from RUL:
Similar works from other Slovenian collections:

Collection

This document is a part of these collections:
  1. Slovenščina 2.0

Back