Details

Stiskanje velikih jezikovnih modelov
ID Burkeljca, Jure (Author), ID Robnik Šikonja, Marko (Mentor) More about this mentor... This link opens in a new window, ID Vreš, Domen (Comentor)

.pdfPDF - Presentation file, Download (579,11 KB)
MD5: 26B193F2433C59B92CE156FFCFC1334D

Abstract
Strukturno obrezovanje velikih jezikovnih modelov (VJM) v kombinaciji z destilacijo znanja in kvantizacijo predstavlja učinkovit pristop za pridobivanje manjših in računsko učinkovitejših velikih jezikovnih modelov. Obrezovanje zmanjša število parametrov oziroma velikost VJM-ja, destilacija znanja omogoča ohranjanje njegove zmogljivosti, kvantizacija pa z uporabo nižje numerične natančnosti dodatno zmanjša porabo pomnilnika in računske zahteve, zlasti med izvajanjem VJM-ja. Preizkusili smo 10 različnih kombinacij obrezovanja v širino in globino ter 2 vrsti kvantizacije, FP8 in NVFP4. Vse različice smo ovrednotili s pomočjo Slovenian LLM Eval, podatke o hitrosti ter učinkovitosti pa smo pridobili s pomočjo knjižnice vLLM. V eksperimentu pokažemo, da je med obravnavanimi pristopi strukturnega obrezovanja v povezavi s kratko destilacijo znanja najučinkovitejše zmanjševanje širine VJM, hkrati pa rezultati potrjujejo visoko učinkovitost kvantizacije v formatu FP8. Na podlagi ugotovitev predlagamo tri pristope za učinkovito stiskanje VJMjev in pripravimo VJM z 8 milijardami parametrov, ki je glede na GaMS3 12B 62,3 % manjši, 84,2 % hitrejši in porabi 53,1 % manj energije, ob 12,8 % nižji kakovosti. Destiliran je bil na približno 2 milijardah žetonov.

Language:Slovenian
Keywords:veliki jezikovni modeli, stiskanje velikih jezikovnih modelov, strukturno obrezovanje velikih jezikovnih modelov, kvantizacija, destilacija znanja, metoda učitelj-učenec
Work type:Bachelor thesis/paper
Organization:FRI - Faculty of Computer and Information Science
Year:2026
PID:20.500.12556/RUL-187808 This link opens in a new window
Publication date in RUL:14.09.2026
Views:14
Downloads:3
Metadata:XML DC-XML DC-RDF
:
Copy citation
Share:Bookmark and Share

Secondary language

Language:English
Title:Large Language Model Compression
Abstract:
Structured pruning of large language models (LLMs) combined with knowledge distillation and quantization represents an effective approach for obtaining smaller and more computationally efficient language models. Pruning reduces the number of parameters and thus the size of the LLM, knowledge distillation helps preserve its performance, while quantization further reduces memory consumption and computational requirements by using lower numerical precision, particularly during LLM inference. We tested ten different combinations of width and depth pruning, as well as two types of quantization, FP8 and NVFP4. We evaluated all variants using Slovenian LLM Eval, while data on speed and efficiency were obtained using the vLLM library. In the experiment, we show that reducing LLM width is the most effective among the evaluated structural compression approaches when combined with short knowledge distillation, while the results also confirm the high efficiency of quantization to the FP8 format. Based on the findings, we propose three efficient LLM compression approaches and develop an 8Bparameter LLM that is 62.3 % smaller, 84.2 % faster, and uses 53.1 % less energy than GaMS3 12B, with a 12.8 % lower quality score after distillation on only 2B tokens.

Keywords:large language models, large language model compression, structured pruning, quantization, knowledge distillation, teacher–student method

Similar documents

Similar works from RUL:
Similar works from other Slovenian collections:

Back