<?xml version="1.0"?>
<rdf:RDF xmlns:rdf="http://www.w3.org/1999/02/22-rdf-syntax-ns#" xmlns:dc="http://purl.org/dc/elements/1.1/"><rdf:Description rdf:about="https://repozitorij.uni-lj.si/IzpisGradiva.php?id=139475"><dc:title>Mining patterns for neurodegenerative diseases from biomedical scientific literature</dc:title><dc:creator>ATANASOSKI,	RADOSLAV	(Avtor)
	</dc:creator><dc:creator>Žitnik,	Slavko	(Mentor)
	</dc:creator><dc:creator>Eftimov,	Tome	(Komentor)
	</dc:creator><dc:subject>data mining</dc:subject><dc:subject>text representation learning</dc:subject><dc:subject>association rule mining</dc:subject><dc:description>Nowadays, there is a vast amount of biomedical knowledge coming in rapidly every day through scientifically published papers. However, trying to keep up with it, is really challenging and takes up too much time. Even more, when searching for relevant papers with required information. To help medical professionals stay up to date, and find papers related to their search topics, in this thesis we create an Information Retrieval (IR) pipeline, first specifying to which neurodegenerative diseases the papers are related to, and also providing analysis to show the most frequent patterns that are researched and published. For the modeling, we explored several state-of-the-art text representation learning models such as BERT, RoBERTa and BioBERT. After fine-tuning each model, BioBERT giving an outstanding performance with 94% cross-validation CA was chosen as a model for the IR pipeline. We also compare our state-of-the-art model with a more traditional and commonly used model, Random Forest. Furthermore, for the analysis of frequent patterns, the abstracts of the diseases involved were annotated and concepts of chemical and genetic compounds were extracted using a Named Entity Recognition (NER) model. After that, all entities were normalized by applying Named Entity Linking (NEL). On the extracted entities, association rule mining was applied in order to find the most frequently researched patterns for each disease, further displayed by using several visualization techniques. These results will help medical professionals to state up to date, on the other side also pointing to missing gaps that are not well researched for a given disease. The data involved in this study was obtained by a publicly available database, PubMed.</dc:description><dc:date>2022</dc:date><dc:date>2022-09-02 14:00:00</dc:date><dc:type>Diplomsko delo/naloga</dc:type><dc:identifier>139475</dc:identifier><dc:language>sl</dc:language></rdf:Description></rdf:RDF>
