Microproteins represent a relatively new and rapidly emerging field in molecular biology and genomics. Advances in analytical and computational technologies have revealed that the protein-coding potential of genomes is substantially greater than previously anticipated. Historically, technical limitations led to the assumption that functional proteins were encoded by open reading frames (ORFs) of at least 300 codons, resulting in the systematic exclusion of short proteins from genomic analyses and annotation. Microproteins are now generally defined as proteins shorter than approximately 100 amino acid residues that are produced through the translation of short ORFs. Despite their small size, microproteins perform important regulatory, signalling, and structural functions and have been identified across all major groups of organisms. Their discovery currently relies on the integration of genomic, translatomic and proteomic approaches supported by advanced bioinformatic methods.
In this master's thesis, we developed a computational pipeline for the processing and analysis of mass spectrometry data acquired using data-dependent acquisition (DDA) and data-independent acquisition (DIA) approaches, employing specialized software MetaMorpheus and DIA-NN. The pipeline was applied to 86 yeast proteomic experiments obtained from the PRIDE database, together with the reference proteome and genome of Saccharomyces cerevisiae. A maximal theoretical proteome library was constructed as an input for the pipeline. To evaluate the resulting proteins, we developed an iterative approach for determining Pareto-optimal fronts, enabling the classification of protein identifications and the definition of criteria for distinguishing credible identifications from false-positive hits. Sequence comparison using BLAST provided additional evidence supporting at least three protein candidates identified from the final candidate set. Two of the highest-ranked candidates were independently confirmed in a related study during the preparation of this thesis and were subsequently incorporated into the latest version of the yeast reference proteome.
The results demonstrate the effectiveness of the developed pipeline for the efficient processing, filtering, and analysis of large-scale mass spectrometry datasets and for the identification of candidate novel proteins. Furthermore, this study highlights the potential of large-scale proteomic data analysis as an alternative to multi-omics approaches for the discovery and characterisation of microproteins.
|