<?xml version="1.0"?>
<metadata xmlns:xsi="http://www.w3.org/2001/XMLSchema-instance" xmlns:dc="http://purl.org/dc/elements/1.1/"><dc:title>Keyword extraction and named entity recognition on Reddit submissions</dc:title><dc:creator>Hudobivnik,	Rok	(Avtor)
	</dc:creator><dc:creator>Helic,	Denis	(Mentor)
	</dc:creator><dc:creator>Bosnić,	Zoran	(Komentor)
	</dc:creator><dc:subject>Globoko učenje</dc:subject><dc:subject>razpoznavanje entitet</dc:subject><dc:subject>luščenje ključnih besed</dc:subject><dc:subject>analiza</dc:subject><dc:description>The goal of this thesis was to create a pipeline for extraction of valuable information from short natural language texts, more specifically Reddit submissions. The two main areas of research that we covered were keyword extraction and named entity recognition for the extraction of keywords and the recognition of actors and movie titles in the texts.  In our thesis we implemented and evaluated four different approaches for keyword extraction (RAKE, TextRank, LSTM and biLSTM networks) and three different approaches for named entity recognition (Spacy library models, Stanford NER and Fine-tuned BERT models). The analysis of the algorithms showed that the best results were achieved when using a three layered biLSTM network for keyword extraction, an uncased BERT model fine-tuned on the MIT movie corpus dataset for the recognition of actors, and the BERT model fine-tuned on the Ontonotes 5 dataset for the recognition of movie titles.</dc:description><dc:date>2020</dc:date><dc:date>2020-07-17 13:10:13</dc:date><dc:type>Magistrsko delo/naloga</dc:type><dc:identifier>117614</dc:identifier><dc:identifier>VisID: 25101</dc:identifier><dc:identifier>COBISS_ID: 17020419</dc:identifier><dc:language>sl</dc:language></metadata>
