This paper describes the process of linguistic error annotation in the Slovene learner corpus KOST 2.0. It presents a starting point for error annotation based on the principle of minimal correction and tending towards the norm of standard Slovene. The error taxonomy is described and illustrated with examples from KOST, in which errors were annotated in 24% of the texts. The taxonomy has 23 types of errors grouped into four basic categories (orthography, vocabulary, morphology, syntax). In practice, the process of error annotation and categorisation has been tested on students of Slovene studies, who annotated 197 corpus texts over the course of three academic years. They found the work interesting and useful, but due to their lack of expertise, lack of experience with Slovene as a foreign language, carelessness and tendency to hypercorrect texts, their results were mostly inadequate.
|