Online media have become a central source of information, yet the volume of their content exceeds the limits of manual analysis and calls for systematic computational approaches. Existing research on the Slovenian media landscape has focused mainly on topic and sentiment analysis, while the structure of relations between actors in media content has received less attention. This master's thesis addresses the question of how to design a system for collecting, cleaning, and analysing online news that enables the study of such relations. To this end, a comprehensive data pipeline was developed, integrating automated news collection via RSS feeds, named entity recognition using the CLASSLA tool, actor normalisation, and co-occurrence network analysis, with results linked to an interactive web interface. Using this system, a corpus of 162,810 articles from 175 registered Slovenian online sources was compiled over a four-month period, with RSS summaries serving as the unit of analysis. The quality of automated entity recognition was assessed through manual validation. The results indicate a markedly uneven distribution of media attention across actors, and a co-occurrence network with a pronounced community structure and small-world characteristics, in which a limited number of actors bridge domestic and international political discourse. The central contribution of the thesis is the developed system itself, that is, a demonstration that such a pipeline can be designed and applied to the Slovenian language, while the presented findings illustrate its usefulness.
|