In this master’s thesis, we present a system based on large language models for affordable and personalized learning support with the aim of reducing learning gaps. The system relies on two large language models: an advanced paid model and a freely available model. The paid model is used to generate an initialization text that is then provided to the free model. For the free model the initialization text defines student’s learning content, the model’s role, and the way explanations should be delivered. The free model then delivers explanations, guides problem solving, and checks understanding within the predefined context.
We evaluated the effectiveness of this approach in an experimental study using mathematics examples. We compared two versions of the free model that received the same input prompt, with one version additionally receiving the initialization text beforehand. We prepared 30 pairs of responses, which were blindly assessed by five independent raters across three dimensions: providing guidance, actionability and revealing the answer. Ratings were given on a scale from -2 to +2; due to the blind setup, raters judged whether the left or the right response was better (much better or slightly better), or whether the difference could not be determined.
The results show a strong advantage for the initialized model, as positive ratings prevail across all three dimensions. For the providing guidance dimension, a small number of negative ratings (favoring the non-initialized model) also appear; such ratings are very rare for actionability and absent for revealing the answer. At the same time, revealing the answer includes a somewhat higher share of ratings indicating that the difference could not be determined. The proportion of strong wins is highest for revealing the answer, followed by actionability, while in providing guidance approximately half of the wins are strong. At the item level, the average differences are clearly positive, confidence intervals remain entirely above zero, and very small p-values indicate that the observed differences are statistically highly convincing. Inter-rater agreement on the 5-point scale is high for revealing the answer and lower for providing guidance and actionability; however, it improves after collapsing the scale to three categories, suggesting that the direction of the outcome was easier to identify than the exact magnitude of the advantage. For actionability, a prevalence effect is also observed: because one category overwhelmingly dominates, chance-corrected agreement coefficients can be lower despite near-unanimous outcomes. We conclude that introducing an initialization text is a meaningful and effective method for improving the quality of learning support and for developing affordable, personalized virtual tutors.
|