Details

Primerjava sistemov za označevanje jezikovnih popravkov v štirih slovenskih besedilnih korpusih
ID Arhar Holdt, Špela (Author), ID Popič, Damjan (Author), ID Stritar Kučuk, Mojca (Author)

.pdfPDF - Presentation file, Download (419,13 KB)
MD5: BBDBA626F63C74688E93515AF06E1AAE
URLURL - Source URL, Visit https://ebooks.uni-lj.si/ZalozbaUL/catalog/book/664/chapter/3901 This link opens in a new window

Abstract
Za slovenščino so na voljo oz. v procesu gradnje štirje besedilni korpusi, ki vsebujejo oznake jezikovnih popravkov: Šolar, KOST, Lektor in STIKit. Popravki v teh korpusih so označeni po različnih označevalnih sistemih, prilagojenih specifikam korpusnega gradiva. Prispevek analizira označevalne sisteme, identificira podobnosti in razlike v označevalnih kategorijah in opredeli možnosti za medsebojne preslikave oznak ter primerjalne analize korpusnega gradiva.

Language:Slovenian
Keywords:slovenščina, besedilni korpusi, jezikovni popravki, KOST, Šolar, Lektor, STIKit
Work type:Article
Typology:1.16 - Independent Scientific Component Part or a Chapter in a Monograph
Organization:FF - Faculty of Arts
FRI - Faculty of Computer and Information Science
Publication status:Published
Publication version:Version of Record
Year:2024
Number of pages:Str. 11-20
PID:20.500.12556/RUL-179610 This link opens in a new window
UDC:811.163.6:004.9
DOI:10.4312/Obdobja.43.11-20 This link opens in a new window
COBISS.SI-ID:215306243 This link opens in a new window
Publication date in RUL:18.02.2026
Views:239
Downloads:103
Metadata:XML DC-XML DC-RDF
:
Copy citation
Share:Bookmark and Share

Record is a part of a monograph

Title:Predpis in norma v jeziku
Editors:Saška Štumberger
Place of publishing:Ljubljana
Publisher:Založba Univerze
Year:2024
ISBN:978-961-297-439-8
COBISS.SI-ID:212955907 This link opens in a new window
Collection title:Zbirka Obdobja
Collection numbering:43
Collection ISSN:1408-211X

Secondary language

Language:English
Abstract:
For Slovenian, four text corpora that contain linguistic error annotations are available or under construction: Šolar, KOST, Lektor, and STIKit. The errors and corrections in these corpora are labeled with different annotation systems, each adapted to the specific characteristics of the corpus material. This article analyses the systems, identifying similarities and differences in the annotation categories, and it explores possibilities for label-mapping and comparative analyses of the corpus material.

Keywords:Slovene, text corpora, error annotation, KOST, Šolar, Lektor, STIKit

Projects

Funder:ARIS - Slovenian Research and Innovation Agency
Project number:J7-3159
Name:Empirična podlaga za digitalno podprt razvoj pisne jezikovne zmožnosti

Funder:ARIS - Slovenian Research and Innovation Agency
Project number:P6-0411
Name:Jezikovni viri in tehnologije za slovenski jezik

Similar documents

Similar works from RUL:
Similar works from other Slovenian collections:

Back