Details

Prilagoditev velikih jezikovnih modelov s človeškimi preferencami
ID Petrič, Timotej (Author), ID Robnik Šikonja, Marko (Mentor) More about this mentor... This link opens in a new window

.pdfPDF - Presentation file, Download (3,23 MB)
MD5: BC66FC185593DB237EE96BEB0E255C58

Abstract
Magistrsko delo naslavlja izboljšanje pogovornih zmožnosti odprtokodnih slovenskih velikih jezikovnih modelov, kot je GaMS 9B Instruct, katerih odgovori so pogosto kratki in slabo strukturirani. Za sistematično zbiranje preferenc slovenskih uporabnikov glede pogovornih zmožnosti modelov smo vzpostavili Slovensko pogovorno areno. Gre za spletno platformo za slepo primerjavo dveh anonimnih modelov, kjer zbrane preference uporabnikov služijo za rangiranje le-teh in zbiranje podatkov za njihovo nadaljnje učenje. Zaradi nezadostne količine zbranih preferenčnih podatkov smo s pomočjo modela GaMS 27B Instruct prevedli visokokakovostno angleško učno množico Nemotron Post Training Dataset v1. Na podlagi te množice smo naučili model GaMS 9B Instruct Nemotron, ki v primerjavi z osnovnim modelom generira daljše in bolje strukturirane odgovore. Uspeh metode potrjujejo visoke ocene uporabnikov v pogovorni areni, kjer se je prilagojeni model uvrstil med najboljše. Model je uspešen tudi na primerjalnih testih Slovenian LLM eval, SloBench in Belebele.

Language:Slovenian
Keywords:veliki jezikovni modeli, Slovenska pogovorna arena, družina modelov GaMS
Work type:Master's thesis/paper
Typology:2.09 - Master's Thesis
Organization:FRI - Faculty of Computer and Information Science
Year:2025
PID:20.500.12556/RUL-173848 This link opens in a new window
COBISS.SI-ID:254192899 This link opens in a new window
Publication date in RUL:24.09.2025
Views:735
Downloads:149
Metadata:XML DC-XML DC-RDF
:
Copy citation
Share:Bookmark and Share

Secondary language

Language:English
Title:Adapting large language models with human preferences
Abstract:
This thesis addresses the challenge of improving the conversational capabilities of open-source Slovene large language models, such as GaMS 9B Instruct, whose responses are often short and poorly structured. To systematically collect user preferences, we established the Slovenian Chatbot Arena, a platform for blind side-by-side comparison of anonymous models where the collected data serves both for model ranking and for gathering training data. Due to the insufficient quantity of collected preference data, we translated the high-quality English dataset Nemotron Post Training Dataset v1 using the GaMS 27B Instruct model. Based on this dataset, we fine-tuned the GaMS 9B Instruct Nemotron model, which, compared to GaMS 9B Instruct, generates longer and better-structured responses. The success of this method is confirmed by high user ratings in the chatbot arena, where the fine-tuned model ranked among the best. The model also performs well on the Slovenian LLM eval, SloBench, and Belebele benchmarks.

Keywords:large language models, Slovenian chatbot arena, GaMS model family

Similar documents

Similar works from RUL:
Similar works from other Slovenian collections:

Back