<?xml version="1.0"?>
<metadata xmlns:xsi="http://www.w3.org/2001/XMLSchema-instance" xmlns:dc="http://purl.org/dc/elements/1.1/"><dc:title>Similarity of arbitrarily long legal documents</dc:title><dc:creator>Vranješ,	Luka	(Avtor)
	</dc:creator><dc:creator>Robnik Šikonja,	Marko	(Mentor)
	</dc:creator><dc:subject>document similarity</dc:subject><dc:subject>document recommendation</dc:subject><dc:subject>legal documents</dc:subject><dc:subject>long documents</dc:subject><dc:subject>natural language processing</dc:subject><dc:subject>transformer neural networks</dc:subject><dc:description>The penetration of modern language technologies into the legal industry is necessary for it to deal with large amounts of texts it produces. Search is a core feature allowing users to perform their work better and faster. The use of modern context-aware approaches can aid in many features related to search, by better quantifying similarity between text.

As a solution, we propose a transformer-based model for creating document embeddings using two interlaced encoders. We train three models with various levels of interlacing and also inform one model of the relative location of each segment within the document. As no differences were detected in the training stage, the most feature rich model was selected and compared in human evaluation to a baseline doc2vec model on a task of recommending similar documents. 

Based on the results, doc2vec proved to be a better and more suitable model for the selected task. The testing outlined some key problems with the proposed model in terms of its concept of similarity, which does not match the requirements of legal document recommendation.</dc:description><dc:date>2022</dc:date><dc:date>2022-10-03 10:25:00</dc:date><dc:type>Magistrsko delo/naloga</dc:type><dc:identifier>141628</dc:identifier><dc:identifier>VisID: 34561</dc:identifier><dc:identifier>COBISS_ID: 125574147</dc:identifier><dc:language>sl</dc:language></metadata>
