<?xml version="1.0"?>
<metadata xmlns:xsi="http://www.w3.org/2001/XMLSchema-instance" xmlns:dc="http://purl.org/dc/elements/1.1/"><dc:title>A case study demonstrating an approach to the statistical analysis of the variation of multiword expressions in Slovene corpora</dc:title><dc:creator>Čibej,	Jaka	(Avtor)
	</dc:creator><dc:subject>multiword expressions</dc:subject><dc:subject>multiword expression variants</dc:subject><dc:subject>statistical analysis</dc:subject><dc:subject>automatic extraction</dc:subject><dc:subject>corpora</dc:subject><dc:description>In Slovene linguistics, much research in phraseology has either been theoretical in nature or focused more on compiling lexicographic resources for human users. While several machine-readable lexicographic resources containing multiword expressions (MWEs) have  also  been  developed  in  recent  years,  Slovene  phraseology  and  computational Slovene linguistics remain largely divided into separate tracks. We attempt to bridge the gap with a brief demonstration of the benefits that computational and statistical approaches based on machine-readable data can have for linguists and phraseologists. We briefly present the SUK Training Corpus of Slovene, the largest machine-readable dataset for Slovene that contains annotations of multiword expressions, as well as the Q-CAT  Corpus  Annotation  Tool that was used to annotate it. We extract examples for two Slovene MWEs (priti  na  zeleno  vejo  and  podirati  se  kot  hišica  iz  kart) from the morphosyntactically annotated Gigafida 2.1 Corpus of Written Standard Slovene using a rule-based approach that leverages syntactic structures. We perform a statistical analysis to determine the degree of variation within the extracted examples. We aim to show that machine-readable data is intended not only for developers of NLP tools but can also help provide additional insight into the structure and variation for the linguistic description of MWEs.</dc:description><dc:date>2025</dc:date><dc:date>2026-01-06 12:28:16</dc:date><dc:type>Drugo</dc:type><dc:identifier>177754</dc:identifier><dc:identifier>UDK: 81'322</dc:identifier><dc:identifier>ISSN pri članku: 0024-3922</dc:identifier><dc:identifier>DOI: 10.4312/linguistica.65.1.45-61</dc:identifier><dc:identifier>COBISS_ID: 263491843</dc:identifier><dc:identifier>OceCobissID: 262399235</dc:identifier><dc:language>sl</dc:language></metadata>
