<?xml version="1.0"?>
<rdf:RDF xmlns:rdf="http://www.w3.org/1999/02/22-rdf-syntax-ns#" xmlns:dc="http://purl.org/dc/elements/1.1/"><rdf:Description rdf:about="https://repozitorij.uni-lj.si/IzpisGradiva.php?id=152353"><dc:title>Deep learning of tissue-specific gene expression from DNA sequences</dc:title><dc:creator>Polanc,	Uroš	(Avtor)
	</dc:creator><dc:creator>Curk,	Tomaž	(Mentor)
	</dc:creator><dc:creator>Zrimec,	Jan	(Komentor)
	</dc:creator><dc:subject>bioinformatics</dc:subject><dc:subject>convolutional neural network</dc:subject><dc:subject>DNA</dc:subject><dc:subject>DNABERT</dc:subject><dc:subject>gene expression</dc:subject><dc:subject>machine learning</dc:subject><dc:subject>sequence motifs</dc:subject><dc:subject>mRNA</dc:subject><dc:subject>predictive models</dc:subject><dc:subject>regulatory mechanisms</dc:subject><dc:subject>tissue-specific gene expression</dc:subject><dc:subject>tissue-specificity</dc:subject><dc:description>Predicting tissue-specific gene expression is a crucial task in understanding the complex regulatory mechanisms governing gene expression. In this research, we employed three distinct models, two convolutional neural networks (CNNs) and DNABERT, to explore predictive models for tissue-specific gene expression. For the genome, we opted for the publicly available \textit{Arabidopsis thaliana}. Our approach involved systematically testing various methodologies, encompassing diverse transcript filtering techniques and an array of input sequences. The integration of multiple models and comprehensive input variations represents a significant step towards enhancing our understanding of tissue-specific gene expression prediction and furthering advancements in bioinformatics and computational biology.

Our findings demonstrate the significance of both sequence data and additional CDS features in predicting gene expression. Combining these features showed only a marginal performance increase. DNABERT struggled with sequence-only inputs but performed comparably to CNN models with augmented CDS features. The Washburn model exhibited the most pronounced tissue-specific performance (R-squared $\approx$ 0.40), followed by DNABERT (R-squared $\approx$ 0.34) and Zrimec (R-squared $\approx$ 0.31). The models faced challenges in predicting both low- and highly-expressed genes but excelled in predicting mid-expressed genes. Additionally, predicting tissue-specific expression closely resembled predicting transcript mean expression, showing a consistent performance ordering across tissues.

We analyzed kernel activations to showcase the model's pattern recognition skills. We cross-referenced these patterns with databases, finding around 650 matches. We used sequence occlusion to pinpoint important areas within the sequences. Our results highlighted the importance of the promoter near the TSS and the 5'UTR near the CDS in shaping model performance, especially with shorter occlusions. Additionally, all genomic regions except the terminator proved relevant when occluding their entire regions.

In conclusion, we have demonstrated the model's capability to forecast tissue-specific gene expression and underscored the significance of non-coding genomic regions. While there remains ongoing research in this field, we aspire that our findings contribute to the understanding of tissue-specific gene expression.</dc:description><dc:date>2023</dc:date><dc:date>2023-11-22 08:45:00</dc:date><dc:type>Magistrsko delo/naloga</dc:type><dc:identifier>152353</dc:identifier><dc:language>sl</dc:language></rdf:Description></rdf:RDF>
