Details

A case study demonstrating an approach to the statistical analysis of the variation of multiword expressions in Slovene corpora
ID Čibej, Jaka (Author)

.pdfPDF - Presentation file, Download (322,56 KB)
MD5: 2C210430AD08DC323B40064B6BED3E70
URLURL - Source URL, Visit https://journals.uni-lj.si/linguistica/article/view/22971 This link opens in a new window

Abstract
In Slovene linguistics, much research in phraseology has either been theoretical in nature or focused more on compiling lexicographic resources for human users. While several machine-readable lexicographic resources containing multiword expressions (MWEs) have also been developed in recent years, Slovene phraseology and computational Slovene linguistics remain largely divided into separate tracks. We attempt to bridge the gap with a brief demonstration of the benefits that computational and statistical approaches based on machine-readable data can have for linguists and phraseologists. We briefly present the SUK Training Corpus of Slovene, the largest machine-readable dataset for Slovene that contains annotations of multiword expressions, as well as the Q-CAT Corpus Annotation Tool that was used to annotate it. We extract examples for two Slovene MWEs (priti na zeleno vejo and podirati se kot hišica iz kart) from the morphosyntactically annotated Gigafida 2.1 Corpus of Written Standard Slovene using a rule-based approach that leverages syntactic structures. We perform a statistical analysis to determine the degree of variation within the extracted examples. We aim to show that machine-readable data is intended not only for developers of NLP tools but can also help provide additional insight into the structure and variation for the linguistic description of MWEs.

Language:English
Keywords:multiword expressions, multiword expression variants, statistical analysis, automatic extraction, corpora
Work type:Other
Typology:1.08 - Published Scientific Conference Contribution
Organization:FRI - Faculty of Computer and Information Science
Publication status:Published
Publication version:Version of Record
Year:2025
Number of pages:Str. 45-61
Numbering:Letn. 65, št. 1
PID:20.500.12556/RUL-177754 This link opens in a new window
UDC:81'322
ISSN on article:0024-3922
DOI:10.4312/linguistica.65.1.45-61 This link opens in a new window
COBISS.SI-ID:263491843 This link opens in a new window
Publication date in RUL:06.01.2026
Views:327
Downloads:145
Metadata:XML DC-XML DC-RDF
:
Copy citation
Share:Bookmark and Share

Record is a part of a proceedings

Title:Frazeologija v digitalni dobi
COBISS.SI-ID:262399235 This link opens in a new window

Record is a part of a journal

Title:Linguistica
Shortened title:Linguistica
Publisher:Filozofska fakulteta, Filozofska fakulteta, Znanstvena založba Filozofske fakultete, Založba Univerze v Ljubljani
ISSN:0024-3922
COBISS.SI-ID:16075264 This link opens in a new window

Licences

License:CC BY-SA 4.0, Creative Commons Attribution-ShareAlike 4.0 International
Link:http://creativecommons.org/licenses/by-sa/4.0/
Description:This Creative Commons license is very similar to the regular Attribution license, but requires the release of all derivative works under this same license.

Secondary language

Language:Slovenian
Title:Študija primera za prikaz pristopa k statistični analizi variantnosti večbesednih enot v slovenskih korpusih
Abstract:
V slovenskem jezikoslovju je precej raziskav na področju frazeologije teoretične na-rave ali pa so usmerjene v gradnjo slovarskih virov za človeške uporabnike. V zadnjih letih je bilo zgrajenih tudi nekaj strojno berljivih slovarskih virov, ki vsebujejo večbe-sedne enote (VBE), a slovenska frazeologija in slovensko računalniško jezikoslovje večinoma ostajata ločeni. V prispevku ju poskusimo povezati s kratkim prikazom ko-risti, ki jih lahko računalniški in statistični pristopi na podlagi strojno berljivih podat-kov prinesejo jezikoslovcem in frazeologom. Na kratko predstavimo učni korpus pisne slovenščine SUK, največjo strojno berljivo slovensko podatkovno množico z oznakami večbesednih enot, ter orodje za označevanje korpusov Q-CAT, s katerim je bil korpus označen. Iz oblikoskladenjsko označenega korpusa pisne standardne slovenščine Giga-fida 2.1 nato izluščimo primere za dve slovenski VBE (priti na zeleno vejo in podirati se kot hišica iz kart) s pomočjo pristopa na podlagi pravil, ki se opira na skladenjske strukture. Izvedemo statistično analizo in določimo stopnjo variantnosti znotraj izluš-čenih primerov. S tem želimo pokazati, da strojno berljivi podatki niso namenjeni le razvijalcem orodij za obdelavo naravnega jezika, temveč lahko tudi jezikoslovcem po-nudijo dodaten vpogled v strukturo in variantnost VBE.

Keywords:večbesedne enote, variante večbesednih enot, statistična analiza, strojno luščenje, korpusi

Similar documents

Similar works from RUL:
Similar works from other Slovenian collections:

Collection

This document is a part of these collections:
  1. Linguistica

Back