Ordinal response variables are common in clinical practice, as they provide more information than binary outcomes. The response variable belongs to naturally ordered categories, such as pain severity or disease stage, where the responses cannot be measured on a numerical scale. When modeling these variables, researchers often face a choice between parsimonious and flexible models. Parsimonious models are based on a smaller number of parameters, stronger structural assumptions, and account for the ordinal structure of the data. Flexible models ignore the order of categories, allowing less restrictive modeling at the cost of increased model complexity. An important methodological dilemma in ordinal regression is the proportional odds (PO) assumption on which the cumulative ordinal regression model is based. In practice, this assumption is often at least partially violated, raising the question of whether an ordinal model remains appropriate or whether a more flexible alternative, such as the multinomial logistic model, should be preferred.
The purpose of this master’s thesis was to assess the impact of PO assumption violations on model predictive performance and to determine under which conditions the cumulative ordinal model remains a suitable choice. The cumulative PO ordinal regression model was compared with the multinomial logistic regression model as a more flexible alternative. Model performance was evaluated through simulation studies as well as an application to real clinical cardiotocography data. In the simulation study, model performance was assessed across different sample sizes (n = 200, 1000, 10000) and three levels of PO assumption violation (none, mild, and severe). Predictive performance was evaluated using the root mean squared error (RMSE) of predicted probabilities, accuracy, and calibration of predictions. In the case study, both models were used to predict an ordinal fetal state (normal, suspect, and pathologic). Modeling was based on cardiotocographic measurement variables. For both models, internal validation and stability analyses were performed. Simulation results showed that the cumulative PO model achieved comparable or better predictive performance and accuracy in most scenarios, particularly for smaller sample sizes. Under strong violations of the PO assumption and large sample sizes, the multinomial model performed better. Calibration results indicated that PO violations did not affect the calibration intercept but influenced the calibration slope, with the multinomial model showing a higher tendency toward overfitting.
In the case study, most predictors did not fully satisfy the PO assumption. Backward elimination based on the BIC criterion in the cumulative model primarily retained variables that at least partially satisfied the assumption, and a similar pattern was also observed for the multinomial model. The selected predictors aligned with factors commonly used in clinical fetal state classification. In the case study, the multinomial model achieved better predictive performance and accuracy, while the cumulative PO model showed a better calibration slope, particularly when predicting pathological outcomes. Overall, the results demonstrate that the cumulative PO model remains a valid and often preferable choice even under moderate violations of the PO assumption, as it preserves good predictive performance while offering improved interpretability and lower model complexity. If sufficiently large samples are available, there are no restrictions on using a more flexible model. However, the question of how large the samples need to be for reliable use remains open.
|