Details

Orkestracija modularnega cevovoda za procesiranje in shranjevanje korpusnih podatkov
ID Sedeljšak, Janez (Author), ID Faganeli Pucer, Jana (Mentor) More about this mentor... This link opens in a new window, ID Kosem, Iztok (Comentor)

.pdfPDF - Presentation file, Download (1,70 MB)
MD5: 2CE2E57527333755D4F8B98CC456A9B4

Abstract
V magistrskem delu se osredotočamo na prenovo podatkovnega cevovoda za obdelavo in shranjevanje medijskih besedil, ki se uporabljajo kot vhod v spremljevalni korpus Trendi. Obstoječi monolitni sistem je podatke hranil izključno v datotečnem sistemu in vmesnih izhodov ni sistematično preverjal, kar je vodilo do pogostih neopaženih napak obdelave in nestabilnosti pri pretvorbi večjih količin besedil. Prenova sledi trem ciljem: zagotoviti preverljivost in sledljivost podatkov, preiti na modularno arhitekturo ter vzpostaviti nadzor, ki izboljša pristop k izvajanju nalog. Cevovod smo razdelili na samostojne komponente. Po ovrednotenju treh podatkovnih baz smo izbrali tisto, ki najbolje ustreza zahtevam sistema in trenutno služi kot operativni vir podatkov cevovoda. Pravilnost podatkov se preverja na prehodih med formati, vsaka zaznana napaka pa izvajanje ustavi in se zabeleži. Bralna nadzorna plošča vzdrževalcem ponuja jasen pregled nad potekom obdelave in trenutnim stanjem podatkov. Pravilnost rešitve smo preverili z integracijskimi testi, osredotočenimi na robne primere, pri čemer smo uporabili referenčne produkcijske podatke.

Language:Slovenian
Keywords:Podatkovni cevovod, Modularnost, Orkestracija, Python, Rust, NLP, Korpus
Work type:Master's thesis/paper
Organization:FRI - Faculty of Computer and Information Science
Year:2026
PID:20.500.12556/RUL-189417 This link opens in a new window
Publication date in RUL:06.10.2026
Views:49
Downloads:8
Metadata:XML DC-XML DC-RDF
:
Copy citation
Share:Bookmark and Share

Secondary language

Language:English
Title:Orchestration of a modular pipeline for processing and storing corpus data
Abstract:
This master's thesis focuses on the redesign of a data pipeline for processing and storing media texts that serve as input to the Trendi monitor corpus. The existing monolithic system stored data exclusively in the file system and did not systematically validate intermediate outputs, which led to frequent unnoticed processing errors and instability when converting larger volumes of text. The redesign pursues three goals: ensuring the verifiability and traceability of data, moving to a modular architecture, and establishing monitoring that improves how tasks are executed. We split the pipeline into independent components. After evaluating three databases, we selected the one that best fits the system's requirements and currently serves as the operational data source of the pipeline. Data correctness is checked at the transitions between formats, and every detected error stops execution and is logged. A read-only dashboard gives maintainers a clear overview of processing progress and the current state of the data. We verified the correctness of the solution with integration tests focused on edge cases, using a reference set of production data.

Keywords:Data pipeline, Modularity, Orchestration, Python, Rust, NLP, Corpus

Similar documents

Similar works from RUL:
Similar works from other Slovenian collections:

Back