<?xml version="1.0"?>
<rdf:RDF xmlns:rdf="http://www.w3.org/1999/02/22-rdf-syntax-ns#" xmlns:dc="http://purl.org/dc/elements/1.1/"><rdf:Description rdf:about="https://repozitorij.uni-lj.si/IzpisGradiva.php?id=167492"><dc:title>Continual learning with superposition in transformers</dc:title><dc:creator>Zeman,	Marko	(Avtor)
	</dc:creator><dc:creator>Bosnić,	Zoran	(Mentor)
	</dc:creator><dc:creator>Faganeli Pucer,	Jana	(Komentor)
	</dc:creator><dc:subject>machine learning</dc:subject><dc:subject>deep learning</dc:subject><dc:subject>continual learning</dc:subject><dc:subject>transformer</dc:subject><dc:subject>superposition</dc:subject><dc:description>The rapid evolution of machine learning and its widespread use across various domains
underscores the imperative for models that learn continuously. Traditional
machine learning models, once trained, remain static, incapable of assimilating a new
task without the risk of catastrophic forgetting, where the acquisition of new knowledge
erases previously learned information. This phenomenon severely limits their
applicability in environments where data and requirements persistently develop.
Addressing this challenge, our dissertation deals with the evolving domain of machine
learning, with a special focus on transformers within the continual learning setting,
marking a path toward achieving computational systems that emulate human adaptability
and learning capabilities. The essence of this research revolves around exploring
and implementing superposition techniques specifically tailored for memory-restrained
devices, such as mobile phones and drones.
The main contribution of our study is the creation of the SuperFormer method,
a novel approach that leverages superposition exclusively during task changes. This
method significantly reduces training time and addresses catastrophic forgetting efficiently,
ensuring optimal use of resources. On a set of NLP classification tasks, Super-
Former achieves the highest AUROC and AUPRC among all comparative methods
while being the fastest to train and needing less additional memory per task than most
of the methods.
Our research goes further than just introducing SuperFormer. It explores how superposition
can be effectively utilized in different fields and with various neural architectures.
We’ve shown that our method is outperforming others also in MLP and CNN
architectures in the computer vision domain.
We also introduced Sparse SuperFormer, which applies sparse learning to boost performance
with fewer weight adjustments, pointing to enhanced model efficiency. Training only half of the weights for each task improved the average accuracy up to 2.2%.
Additionally, we developed the SuperAdapter strategy to increase memory efficiency
in continual learning. By combining SuperFormer with adapters, it’s possible to learn
multiple tasks within one adapter, minimizing storage impact with minimal loss in performance.
Looking at the average AUPRC, the adapter’s storage requirements can be
halved by losing only 2.0 to 5.6%, depending on the adapter size.
In conclusion, this dissertation represents a significant advancement in our understanding
and application of continual learning and superposition. As we look towards
the future, these advancements have the potential to significantly impact various fields
by enabling AI systems to learn and evolve in dynamic environments.</dc:description><dc:date>2025</dc:date><dc:date>2025-02-24 14:52:56</dc:date><dc:type>Doktorsko delo/naloga</dc:type><dc:identifier>167492</dc:identifier><dc:language>sl</dc:language></rdf:Description></rdf:RDF>
