<?xml version="1.0"?>
<rdf:RDF xmlns:rdf="http://www.w3.org/1999/02/22-rdf-syntax-ns#" xmlns:dc="http://purl.org/dc/elements/1.1/"><rdf:Description rdf:about="https://repozitorij.uni-lj.si/IzpisGradiva.php?id=146738"><dc:title>Topic analysis of Slovenian news and social media</dc:title><dc:creator>HLADNIK,	JUŠ	(Avtor)
	</dc:creator><dc:creator>Robnik Šikonja,	Marko	(Mentor)
	</dc:creator><dc:subject>topic modeling</dc:subject><dc:subject>language models</dc:subject><dc:subject>Slovene language</dc:subject><dc:subject>topic model stability and similarity</dc:subject><dc:subject>natural language processing</dc:subject><dc:description>Topic modeling is an unsupervised machine learning technique that aims to discover hidden semantic structures within large collections of text documents, thus facilitating the exploration and understanding of vast textual data.
We conduct a comprehensive comparison of four popular topic modeling algorithms, namely LDA, NMF, Top2vec and BERTopic, in the context of the Slovenian language. To assess the performance of these algorithms we use topic coherence and topic diversity quantitative evaluation and additionally manually interpret extracted topics. Our results demonstrate that all models achieve higher topic coherence on the news corpus compared to tweets. While BERTopic is the only algorithm to produce satisfactory results on the tweets corpus, all models perform well on the news corpus.
Furthermore, we introduce a novel method, MBTS (Maximum Bipartite Topic Similarity), for comparing the similarity of topic models and evaluating their stability. This method relies on semantic similarity and maximum graph bipartite matching. Our findings have important implications for the selection and application of topic modeling algorithms in the context of the Slovenian language. Moreover, the MBTS method opens up a new and important area of topic model stability evaluation.</dc:description><dc:date>2023</dc:date><dc:date>2023-06-09 15:10:00</dc:date><dc:type>Magistrsko delo/naloga</dc:type><dc:identifier>146738</dc:identifier><dc:language>sl</dc:language></rdf:Description></rdf:RDF>
