Predicting future values in longitudinal data is challenging because repeated measurements from the same individual are correlated. In addition, future values of time-varying explanatory variables are usually not known at the time of prediction. In this master's thesis, we compared linear mixed-effects models and machine learning methods and examined how their predictive performance is affected by the size of the training set, model misspecification, and the way previous measurements are used for prediction.
The comparison was based on a simulation study with 100 Monte Carlo replications and four simulation scenarios. In each replication, a population of 1,000 individuals with 12 measurements per individual was generated. The data were balanced and contained no missing values. Three nested training sets with 50, 100, and 250 individuals were created from the same population, while the remaining 750 individuals formed a common test set. For individuals in the test set, the first eight measurements were known and the final four were predicted. We compared a linear model, several versions of a linear mixed-effects model, a landmark LMM, random forest, and XGBoost. Predictions of the final four measurements were evaluated using RMSE, MAE, and bias. For the classical LMM, the specification of the fixed and random parts of the model was also assessed using a diagnostic test.
The results did not show one model that was best in all settings. The best method depended on the simulation scenario and the size of the training set. With smaller training sets, the classical or scenario-adapted LMM generally performed best. With 250 training individuals, the landmark LMM had the lowest mean RMSE and MAE in all four scenarios. Misspecification of the linear mixed-effects model generally reduced predictive performance. The diagnostic procedure was able to detect problems in both the fixed and random parts of the model. For random forest and XGBoost, using more features did not necessarily lead to better predictions. Their performance depended on how well the input features represented the temporal structure of the data.
The results show that the choice of model for longitudinal prediction should depend on the structure of the data, the sample size, and the available measurement history. Goodness-of-fit diagnostics and evaluation on new individuals provide different but complementary information: the diagnostic can indicate possible problems with the model specification, while prediction on new individuals shows whether changing the model actually improves the predictions.
|