<?xml version="1.0"?>
<metadata xmlns:xsi="http://www.w3.org/2001/XMLSchema-instance" xmlns:dc="http://purl.org/dc/elements/1.1/"><dc:title>Scalable matrix factorization for data fusion</dc:title><dc:creator>Čopar,	Andrej	(Avtor)
	</dc:creator><dc:creator>Zupan,	Blaž	(Mentor)
	</dc:creator><dc:subject>Machine learning</dc:subject><dc:subject>bioinformatics</dc:subject><dc:subject>matrix factorization</dc:subject><dc:subject>data fusion</dc:subject><dc:description>Data collection technologies are advancing quickly and are producing larger amounts of data than ever before. Biomedical data analysis, text analysis and recommender systems rely on machine learning to perform tasks such as modeling gene-disease associations, clustering documents, and user recommendations. The analysis of such data is particularly challenging due to the large dimensionality and multitude of different object types. Data fusion methods can accurately deal with such heterogeneous datasets by integrating them into a single model. Existing data fusion approaches were not designed for speed on huge datasets and can be prohibitively slow for practical use. Our main goal is to develop new methods that increase the speed of data fusion using efficient optimization techniques and modern parallel systems. 

Contemporary data fusion methods are based on matrix factorization as its core component. Matrix factorization learns a latent data model that transforms the data into a latent feature space enabling generalization, noise removal and feature discovery. Matrix tri-factorization is a popular method that is not limited by the assumption of standard matrix factorization about data residing in one latent space.  Matrix tri-factorization infers separate latent space for each dimension, making the approach ideal for data fusion. Factorization algorithms are numerically intensive, hence scaling current algorithms to work with large datasets is crucial for development of fast data fusion approaches.

We developed a block\-/wise approach for latent factor learning in matrix tri\-/factorization. The approach partitions a data matrix into disjoint submatrices that are treated independently and fed into a parallel factorization system. We show that our approach scales well on multi-processor and multi-GPU architectures. Our approach on four GPU devices is more than a hundred times faster than its single-processor counterpart.

Currently, non-negative matrix tri-factorization learns a representation of a dataset through an optimization procedure that typically uses multiplicative update rules. This procedure has had limited success due to its slow convergence. We develop three alternative optimization techniques for non-negative matrix tri-factorization based on alternating least squares, projected gradients, and coordinate descent. We perform an empirical study comparing multiplicative update rules with the three alternative techniques and show that coordinate descent-based techniques converges up to twenty times faster compared to multiplicative updates.

Finally, we employ block-wise techniques together with coordinate descent to speed up data fusion. With block-wise parallelization we accelerate an existing data fusion approach over 30 times. We derive a new coordinate descent-based data fusion approach that converges over 15 times faster compared to existing approach. Coordinate-descent data fusion accelerated on GPU devices performs over 100 times faster compared to an existing approach on 16 processes.</dc:description><dc:date>2019</dc:date><dc:date>2019-11-19 12:11:45</dc:date><dc:type>Doktorsko delo/naloga</dc:type><dc:identifier>112894</dc:identifier><dc:identifier>VisID: 19682</dc:identifier><dc:identifier>COBISS_ID: 1538450627</dc:identifier><dc:language>sl</dc:language></metadata>
