This thesis addresses the challenge of improving the conversational capabilities of open-source Slovene large language models, such as GaMS 9B Instruct, whose responses are often short and poorly structured. To systematically collect user preferences, we established the Slovenian Chatbot Arena, a platform for blind side-by-side comparison of anonymous models where the collected data serves both for model ranking and for gathering training data. Due to the insufficient quantity of collected preference data, we translated the high-quality English dataset Nemotron Post Training Dataset v1 using the GaMS 27B Instruct model. Based on this dataset, we fine-tuned the GaMS 9B Instruct Nemotron model, which, compared to GaMS 9B Instruct, generates longer and better-structured responses. The success of this method is confirmed by high user ratings in the chatbot arena, where the fine-tuned model ranked among the best. The model also performs well on the Slovenian LLM eval, SloBench, and Belebele benchmarks.
|