<?xml version="1.0"?>
<rdf:RDF xmlns:rdf="http://www.w3.org/1999/02/22-rdf-syntax-ns#" xmlns:dc="http://purl.org/dc/elements/1.1/"><rdf:Description rdf:about="https://repozitorij.uni-lj.si/IzpisGradiva.php?id=104021"><dc:title>Low-rank matrix factorization in multiple kernel learning</dc:title><dc:creator>Stražar,	Martin	(Avtor)
	</dc:creator><dc:creator>Curk,	Tomaž	(Mentor)
	</dc:creator><dc:subject>Strojno učenje</dc:subject><dc:subject>bioinformatika</dc:subject><dc:subject>matrična faktorizacije</dc:subject><dc:subject>jerdne metode</dc:subject><dc:subject>učenje z več jedrnimi funkcijami</dc:subject><dc:subject>linearna regresija</dc:subject><dc:subject>interakcije proteini-RNA</dc:subject><dc:description>The increased rate of data collection, storage, and availability results in a corresponding interest for data analyses and predictive models based on simultaneous inclusion of multiple data sources. This tendency is ubiquitous in practical applications of machine learning, including recommender systems, social network analysis, finance and computational biology. The heterogeneity and size of the typical datasets calls for simultaneous dimensionality reduction and inference from multiple data sources in a single model. Matrix factorization and multiple kernel learning models are two general approaches that satisfy this goal. This work focuses on two specific goals, namely i) finding interpretable, non-overlapping (orthogonal) data representations through matrix factorization and ii) regression with multiple kernels through the low-rank approximation of the corresponding kernel matrices, providing non-linear outputs and interpretation of kernel selection.  

The motivation for the models and algorithms designed in this work stems from RNA biology and the rich complexity of protein-RNA interactions. Although the regulation of RNA fate happens at many levels - bringing in various possible data views - we show how different questions can be answered directly through constraints in the model design. We have developed an integrative orthogonality nonnegative matrix factorization (iONMF) to integrate multiple data sources and discover non-overlapping, class-specific RNA binding patterns of varying strengths. We show that the integration of multiple data sources improves the predictive accuracy of retrieval of RNA binding sites and report on a number of inferred protein-specific patterns, consistent with experimentally determined properties. 
 
A principled way to extend the linear models to non-linear settings are kernel methods.  Multiple kernel learning enables modelling with different data views, but are limited by the quadratic computation and storage complexity of the kernel matrix.  Considerable savings in time and memory can be expected if kernel approximation and multiple kernel learning are performed simultaneously.  We present the Mklaren algorithm, which achieves this goal via Incomplete Cholesky Decomposition, where the selection of basis functions is based on Least-angle regression, resulting in linear complexity both in the number of data points and kernels.  Considerable savings in approximation rank are observed when compared to general kernel matrix decompositions and comparable to methods specialized to particular kernel function families. The principal advantages of Mklaren are independence of kernel function form, robust inducing point selection and the ability to use different kernels in different regions of both continuous and discrete input spaces, such as numeric vector spaces, strings or trees, providing a platform for bioinformatics.  

In summary, we design novel models and algorithms based on matrix factorization and kernel learning, combining regression, insights into the domain of interest by identifying relevant patterns, kernels and inducing points, while scaling to millions of data points and data views.</dc:description><dc:date>2018</dc:date><dc:date>2018-10-01 08:40:07</dc:date><dc:type>Doktorsko delo/naloga</dc:type><dc:identifier>104021</dc:identifier><dc:language>sl</dc:language></rdf:Description></rdf:RDF>
