<?xml version="1.0"?>
<rdf:RDF xmlns:rdf="http://www.w3.org/1999/02/22-rdf-syntax-ns#" xmlns:dc="http://purl.org/dc/elements/1.1/"><rdf:Description rdf:about="https://repozitorij.uni-lj.si/IzpisGradiva.php?id=102887"><dc:title>Improvement of data dependent acquisition analysis in mass spectrometry</dc:title><dc:creator>BAUMKIRCHER,	ALJAŽ	(Avtor)
	</dc:creator><dc:creator>Reberšek,	Matej	(Mentor)
	</dc:creator><dc:subject>Pulsar algorithm</dc:subject><dc:subject>data-dependent acquisition</dc:subject><dc:subject>search engine</dc:subject><dc:subject>machine learning</dc:subject><dc:subject>proteomics</dc:subject><dc:description>The field of proteomics is a constantly evolving scientific field, focused on the large-scale study of proteins. In general, researchers can choose between targeted analysis, where they observe a specific protein, or discovery analysis, where they get a qualitative and possibly quantitative overview of the whole proteome. There are different approaches to the discovery analysis, with the data-dependent acquisition (DDA) currently being one of the most widely adopted analysis approaches in the field of proteomics. 

This thesis focuses on the Pulsar algorithm, a search engine developed by Biognosys AG (Switzerland). Although Pulsar algorithm is able to analyze different types of acquisitions (DDA, DIA, PRM), we will focus exclusively on DDA data acquisition. The overall quality of a search engine is determined by the amount of identified proteins, as well as the execution time needed for the analysis. The goal of the thesis is to provide a general overview on how the Pulsar algorithm works, as well as how well it performs compared to some other search engines on the market (e.g. MaxQuant, SEQUEST integrated in the Thermo Proteome Discoverer). Furthermore, we investigate a possibility of improving the execution time of the analysis by optimizing the calibration process. Lastly, we tried to evaluate how a machine learning tool (e.g. Percolator) could bring additional value to the Pulsar algorithm.

With approaches described in the chapter Materials and Methods and results presented and evaluated in the chapter Results and Discussions we managed to achieve a 11.06% of an improvement regarding the analysis execution time (Approach 3.3, Table 3.6), while not significantly influencing the number of identified protein groups (Approach 3.3, Table 3.5). Moreover, implementing Percolator in the Pulsar algorithm's workflow results in an average increase of 20.20% in protein group identifications, while consuming only 5.61% of the total execution time (Table 3.7). Combining both methods could potentially result in around 20% of an improvement in the protein group identifications, without prolonging the execution time. 

Since SEQUEST (using the Percolator) has on average 17.35% more protein group identifications than the Pulsar algorithm (Table 3.9), one could argue that with the modifications described in this thesis the Pulsar algorithm could be comparable, if not even better, than all other search engines used in this thesis.

In the end, the reader should have a good overview of the field of mass spectrometry, data-dependent acquisition, and the quality of the Pulsar algorithm in comparison to other well-established search engines on the market.</dc:description><dc:date>2018</dc:date><dc:date>2018-09-11 14:55:02</dc:date><dc:type>Magistrsko delo/naloga</dc:type><dc:identifier>102887</dc:identifier><dc:language>sl</dc:language></rdf:Description></rdf:RDF>
