In this paper, we present the process of compiling an empirically based comprehensive list of accentuated units in Slovene based on the Sloleks Morphological Lexicon of Slovene, with an emphasis on constantly accentuated units. Existing language manuals (e.g. the Slovene Grammar and the Slovenian Normative Guide 2001) provide only partial data without specified frequencies. A machine-readable list of constantly accentuated units would be useful not only to language speakers, but also for automatic accentuators due to the only partly predictable place of accentuation in Slovene. We first export all accentuated units from the lexicon, then for each of the exported accentuated word parts we count how many lexemes it occurs in. We calculate the proportions in which the relevant word part is accentuated. We analyse the data obtained in this way by word types and compare the differences between our results and existing Slovene language manuals. We describe the structure of the resulting list, which is openly accessible in the CLARIN.SI repository, and document the process of its compilation. Using the same process, a new version of the list can always be exported when the lexicon is updated, taking into account new data in order to accurately reflect accentuation in Slovene. We finish the paper by listing several possibilities for improvement and potential steps for further research.
|