Details

Primerjava mešanih modelov in modelov strojnega učenja za napovedovanje longitudinalnih podatkov
ID Ratajec, Mariša (Author), ID Blagus, Rok (Mentor) More about this mentor... This link opens in a new window

.pdfPDF - Presentation file, Download (1,18 MB)
MD5: CF9023C7745A8C95A752E4FFA0C05F4D

Abstract
Napovedovanje prihodnjih vrednosti pri longitudinalnih podatkih je zahtevno, saj so ponavljajoče se meritve istega posameznika med seboj povezane. Poleg tega ob času napovedovanja prihodnje vrednosti časovno odvisnih pojasnjevalnih spremenljivk praviloma še niso znane. V magistrskem delu smo primerjali linearne mešane modele in metode strojnega učenja ter preverili, kako na njihovo napovedno uspešnost vplivajo velikost učne množice, napačna specifikacija modela in način uporabe preteklih meritev posameznika. Primerjava je temeljila na simulacijski študiji s 100 Monte Carlo ponovitvami in štirimi scenariji generiranja podatkov. V vsaki ponovitvi je bilo ustvarjenih 1000 oseb z 12 meritvami. Podatki so bili uravnoteženi in brez manjkajočih vrednosti. Iz iste populacije so bile oblikovane ugnezdene učne množice s 50, 100 in 250 osebami, preostalih 750 oseb pa je tvorilo skupno testno množico. Pri testnih osebah je bilo prvih osem meritev znanih, napovedovali pa smo zadnje štiri. Primerjali smo linearni model, več različic linearnega mešanega modela (LMM), LMM z mejnikom, naključni gozd in XGBoost. Napovedi zadnjih štirih meritev smo ovrednotili z RMSE, MAE in pristranskostjo. Pri klasičnem LMM smo dodatno preverjali ustreznost fiksnega in naključnega dela modela z diagnostičnim testom. Rezultati niso pokazali enega najboljšega modela za vse okoliščine. Najboljši pristop je bil odvisen od scenarija in velikosti učne množice. Pri manjših učnih množicah sta bila praviloma najuspešnejša klasični ali scenarijsko prilagojeni LMM, pri 250 učnih osebah pa je LMM z mejnikom dosegel najnižji povprečni RMSE in MAE v vseh štirih scenarijih. Napačna specifikacija linearnega mešanega modela je praviloma poslabšala napovedno uspešnost. Uporabljena diagnostika je uspešno zaznavala odstopanja v fiksnem in naključnem delu modela. Pri naključnem gozdu in XGBoostu večje število značilk samo po sebi ni zagotavljalo boljše napovedi. Uspešnost je bila odvisna od tega, kako dobro je vhodna predstavitev odražala časovno strukturo podatkov. Ugotovitve kažejo, da je treba model za longitudinalno napovedovanje izbrati glede na strukturo podatkov, velikost učne množice in razpoložljivo zgodovino. Diagnostika prileganja in vrednotenje na novih osebah se pri tem dopolnjujeta: prva opozori na možno strukturno neskladje, drugo pa pokaže, ali sprememba modela dejansko izboljša napovedi.

Language:Slovenian
Keywords:longitudinalni podatki, linearni mešani modeli, analiza z mejnikom, strojno učenje, napovedovanje, simulacijska študija, diagnostika prileganja
Work type:Master's thesis/paper
Typology:2.09 - Master's Thesis
Organization:FE - Faculty of Electrical Engineering
Year:2026
PID:20.500.12556/RUL-186608 This link opens in a new window
COBISS.SI-ID:290461443 This link opens in a new window
Publication date in RUL:03.09.2026
Views:233
Downloads:39
Metadata:XML DC-XML DC-RDF
:
Copy citation
Share:Bookmark and Share

Secondary language

Language:English
Title:A comparison of mixed-effects models and machine learning methods for longitudinal prediction
Abstract:
Predicting future values in longitudinal data is challenging because repeated measurements from the same individual are correlated. In addition, future values of time-varying explanatory variables are usually not known at the time of prediction. In this master's thesis, we compared linear mixed-effects models and machine learning methods and examined how their predictive performance is affected by the size of the training set, model misspecification, and the way previous measurements are used for prediction. The comparison was based on a simulation study with 100 Monte Carlo replications and four simulation scenarios. In each replication, a population of 1,000 individuals with 12 measurements per individual was generated. The data were balanced and contained no missing values. Three nested training sets with 50, 100, and 250 individuals were created from the same population, while the remaining 750 individuals formed a common test set. For individuals in the test set, the first eight measurements were known and the final four were predicted. We compared a linear model, several versions of a linear mixed-effects model, a landmark LMM, random forest, and XGBoost. Predictions of the final four measurements were evaluated using RMSE, MAE, and bias. For the classical LMM, the specification of the fixed and random parts of the model was also assessed using a diagnostic test. The results did not show one model that was best in all settings. The best method depended on the simulation scenario and the size of the training set. With smaller training sets, the classical or scenario-adapted LMM generally performed best. With 250 training individuals, the landmark LMM had the lowest mean RMSE and MAE in all four scenarios. Misspecification of the linear mixed-effects model generally reduced predictive performance. The diagnostic procedure was able to detect problems in both the fixed and random parts of the model. For random forest and XGBoost, using more features did not necessarily lead to better predictions. Their performance depended on how well the input features represented the temporal structure of the data. The results show that the choice of model for longitudinal prediction should depend on the structure of the data, the sample size, and the available measurement history. Goodness-of-fit diagnostics and evaluation on new individuals provide different but complementary information: the diagnostic can indicate possible problems with the model specification, while prediction on new individuals shows whether changing the model actually improves the predictions.

Keywords:longitudinal data, linear mixed-effects models, landmark approach, machine learning, prediction, simulation study, goodness-of-fit diagnostics

Similar documents

Similar works from RUL:
Similar works from other Slovenian collections:

Back