Details

Neuroevolucija adapterjev za velike jezikovne modele
ID Perman, Katarina (Author), ID Robnik Šikonja, Marko (Mentor) More about this mentor... This link opens in a new window

.pdfPDF - Presentation file, Download (3,24 MB)
MD5: 0638DFDDF58A940FE70DE3AB3242B0B6

Abstract
Naloga obravnava problem postavljanja adapterjev v velike jezikovne modele za doseganje čim boljših rezultatov. Dandanes so veliki jezikovni modeli ključno orodje za procesiranje naravnega jezika. Uporabljamo jih lahko za vrsto nalog, a jih moramo za te naloge najprej prilagoditi. Učenje celotnega modela, ki je pogosto zelo velik, je računsko drago in časovno zamudno, zato namesto tega uporabljamo adapterje, ki jih moramo postaviti na prava mesta v zgradbi velikega jezikovnega modela in jim izbrati pravilne hiperparametre, da so uspešni. Problema se lotimo z uporabo genetskih algoritmov, saj bi za preiskavo vseh možnih konfiguracij potrebovali preveč časa. S tem ne pridobimo le konkretnih rešitev, ampak tudi vpogled v vzorce konfiguracij, ki se pojavljajo v boljših rešitvah. To smo preizkusili na modelu BERT z adapterjem LoRA za klasifikacijo čustev. Ugotovimo, da je večanje ranga v adapterju bolj učinkovito kot spreminjanje postavitve, a večji rang pomeni več parametrov, zato moramo ob tem biti previdni. Naši rezultati kažejo tudi, da je v našem primeru največja smiselna količina parametrov adapterja do 5% količine parametrov originalnega modela, saj več parametrov ne pripomore več k višji natančnosti, in da je najboljša možna natančnost za neko velikost adapterja omejena navzgor.

Language:Slovenian
Keywords:velik jezikovni model, adapter, neuroevolucija, genetski algoritmi
Work type:Bachelor thesis/paper
Typology:2.11 - Undergraduate Thesis
Organization:FRI - Faculty of Computer and Information Science
Year:2025
PID:20.500.12556/RUL-173303 This link opens in a new window
COBISS.SI-ID:253637379 This link opens in a new window
Publication date in RUL:15.09.2025
Views:480
Downloads:172
Metadata:XML DC-XML DC-RDF
:
Copy citation
Share:Bookmark and Share

Secondary language

Language:English
Title:Neuroevolution of adapters for large language models
Abstract:
This work touches upon the problem of where to put adapters in large lan- guage models to achieve the best performance possible. Large language - models are nowadays a key tool for natural language processing. We use them for an array of tasks, but we must first fine-tune them for these tasks. Training the whole model, which is often very large, is computationally ex- pensive and time consuming, so we use adapters instead, but we must place them in the right places in the large language model architecture and pick the right hyperparameters for them to be successful. We approach this problem by making use of genetic algorithms, as checking every possible configura- tion would take way too long. We tested this on the BERT model with the LoRA adapter, on emotion classficiation. That way we get not only con- crete solutions but also insight into the patterns that appear in favorable solutions. We find that increasing the rank of the adapters is more effective than changing their placement, but a higher rank means more parameters, so care must be taken. Our results also show that the biggest sensible amount of parameters is up to 5% the amount of parameters of the original model, as more parameters don’t contribute to higher accuracy, and that the best possible accuracy for a certain adapter size has an upper bound.

Keywords:large language model, adapter, neuroevolution, genetic algorithms

Similar documents

Similar works from RUL:
Similar works from other Slovenian collections:

Back