<?xml version="1.0"?>
<metadata xmlns:xsi="http://www.w3.org/2001/XMLSchema-instance" xmlns:dc="http://purl.org/dc/elements/1.1/"><dc:title>Data embedding and fusion by tropical matrix factorization</dc:title><dc:creator>Omanović,	Amra	(Avtor)
	</dc:creator><dc:creator>Oblak,	Polona	(Mentor)
	</dc:creator><dc:creator>Curk,	Tomaž	(Komentor)
	</dc:creator><dc:subject>data mining</dc:subject><dc:subject>data embedding</dc:subject><dc:subject>matrix factorization</dc:subject><dc:subject>tropical factorization</dc:subject><dc:subject>subtropical semiring</dc:subject><dc:subject>tropical semiring</dc:subject><dc:subject>sparse data</dc:subject><dc:subject>matrix completion</dc:subject><dc:description>Data embedding and fusion represent one of the main challenges in machine learning. Meaningful low-dimensional representations of real-world data help algorithms to perform different data mining and prediction tasks successfully. Matrix factorization methods embed data into a latent space using a two-factorization or tri-factorization approaches. These methods mostly use standard linear algebra, which is limited in modeling complex patterns. The non-linearity can be modeled by using tropical semiring, which enables a better approximation of extreme values and distributions, thus discovering high-variance patterns that differ from those found by standard linear algebra.
The motivation for creating data embedding and fusion methods by tropical matrix factorization is found in properties such as non-linearity, the ability to interpret results easily, the intuition behind path-finding problems in graphs, connections with neural networks, and the lack of tropical methods in data mining and machine learning. 

In the thesis, we design novel models and algorithms for data embedding and fusion based on tropical matrix factorization with theoretical and experimental evaluation.
We have developed a sparse tropical matrix factorization (STMF), which returns two factor matrices and performs matrix completion. We apply STMF to predict gene expression values on multiple TCGA datasets. We show that STMF expresses extreme values very well and is robust to overfitting. The main drawback of STMF is slow computational performance, so we propose an efficient version of STMF called FastSTMF. Results showed that FastSTMF outperforms STMF by achieving higher performance, such as faster convergence speed and better approximation results.

In data fusion, tri-factorization methods achieve superior results than two-factorization by utilizing an intermediate approach for fusion of multiple data sources. We present the tropical matrix tri-factorization algorithm called triFastSTMF, which we apply to recover the edge lengths of a four-partition network. We use triFastSTMF to create a tropical data fusion method (tropDF) and show its correctness and convergence through experimental evaluation.</dc:description><dc:date>2023</dc:date><dc:date>2023-10-25 15:50:01</dc:date><dc:type>Doktorsko delo/naloga</dc:type><dc:identifier>151928</dc:identifier><dc:identifier>VisID: 28001</dc:identifier><dc:identifier>COBISS_ID: 172011011</dc:identifier><dc:language>sl</dc:language></metadata>
