This master's thesis focuses on the redesign of a data pipeline for processing and storing media texts that serve as input to the Trendi monitor corpus. The existing monolithic system stored data exclusively in the file system and did not systematically validate intermediate outputs, which led to frequent unnoticed processing errors and instability when converting larger volumes of text. The redesign pursues three goals: ensuring the verifiability and traceability of data, moving to a modular architecture, and establishing monitoring that improves how tasks are executed. We split the pipeline into independent components. After evaluating three databases, we selected the one that best fits the system's requirements and currently serves as the operational data source of the pipeline. Data correctness is checked at the transitions between formats, and every detected error stops execution and is logged. A read-only dashboard gives maintainers a clear overview of processing progress and the current state of the data. We verified the correctness of the solution with integration tests focused on edge cases, using a reference set of production data.
|