Drinking water quality is a key public health factor and its monitoring is mandated by law. In Slovenia, a regular national monitoring programme tracks a wide range of microbiological and chemical quality indicators. With the growing volume of measurements collected, the question arises whether machine learning methods can improve or complement the traditional compliance assessment. The goal of this thesis is to compare a rule-based reference approach and machine learning methods for predicting the compliance of drinking water samples with prescribed quality indicators. The task is formulated as a binary classification problem on a large dataset of samples from the Slovenian national monitoring programme for the period 2019-2025. Approaches are evaluated using Leave-One-Year-Out time-series cross-validation, with feature selection performed using the SHAP method. We compare the rule-based reference approach against four machine learning models: Gradient Boosting, Random Forest, LightGBM, and a Multi-Layer Perceptron. Two novel approaches are also proposed: a regulatory correction for metolachlor aimed at the correct treatment of degradation products (ESA, OXA) in the pesticide sum assessment, and a two-step model for improved handling of rare violations. Results show that machine learning models achieve comparable performance to the reference approach. The proposed regulatory correction for metolachlor substantially reduces the number of errors in the reference approach.
|