For Slovenian, four text corpora that contain linguistic error annotations are available or under construction: Šolar, KOST, Lektor, and STIKit. The errors and corrections in these corpora are labeled with different annotation systems, each adapted to the specific characteristics of the corpus material. This article analyses the systems, identifying similarities and differences in the annotation categories, and it explores possibilities for label-mapping and comparative analyses of the corpus material.
|