<?xml version="1.0"?>
<rdf:RDF xmlns:rdf="http://www.w3.org/1999/02/22-rdf-syntax-ns#" xmlns:dc="http://purl.org/dc/elements/1.1/"><rdf:Description rdf:about="https://repozitorij.uni-lj.si/IzpisGradiva.php?id=185098"><dc:title>On simple EM acceleration schemes suitable for mixture modelling with high overlap between components</dc:title><dc:creator>Panić,	Branislav	(Avtor)
	</dc:creator><dc:creator>Klemenc,	Jernej	(Avtor)
	</dc:creator><dc:creator>Nagode,	Marko	(Avtor)
	</dc:creator><dc:creator>Oman,	Simon	(Avtor)
	</dc:creator><dc:subject>mixture modelling</dc:subject><dc:subject>expectation-maximisation</dc:subject><dc:subject>acceleration</dc:subject><dc:subject>parameter estimation</dc:subject><dc:description>The Expectation-Maximisation (EM) algorithm is widely used for maximum likelihood estimation in incomplete data problems such as mixture modelling, but it often converges slowly, particularly when mixture components overlap substantially. This study presents a comprehensive empirical evaluation of simple EM acceleration schemes for Gaussian mixture models, comparing linear (STEM), quadratic (SQUAREM), and greedy (line search, golden section) methods across 240 simulated mixture configurations spanning three dimensionalities, four component counts, five overlap levels, and four sample sizes. A key contribution is the first systematic comparison of the three acceleration parameter estimates (▫$\alpha$▫▫$_1$▫, ▫$\alpha$▫▫$_2$▫, ▫$\alpha$▫▫$_3$▫) in the mixture modelling context: we show that only ▫$\alpha$▫▫$_3$▫, which is derived as the geometric mean estimate of ▫$\alpha$▫▫$_1$▫ and ▫$\alpha$▫▫$_2$▫, provides genuine acceleration, while ▫$\alpha$▫▫$_1$▫ and ▫$\alpha$▫▫$_2$▫ consistently increase iteration counts by 50–110% relative to ▫$\alpha$▫▫$_3$▫, effectively acting as deceleration. With ▫$\alpha$▫▫$_3$▫, SQUAREM reduces iterations by up to 48% with negligible computational overhead, while greedy methods achieve similar iteration reductions but at 50–110% greater wall-clock time due to repeated log-likelihood evaluations. Crucially, acceleration does not degrade parameter estimation quality under any tested combination of initialisation, overlap, dimensionality, or number of components. We further examine the interaction between acceleration and initialisation, finding that k-means benefits most from acceleration (up to 50% time savings), while the REBMIX (Rough-Enhanced-Bayes MIXture estimation) algorithm benefits least as it already starts near the optimum. Among REBMIX configurations, histogram preprocessing with the outliers mode traversing strategy offers the best trade-off between quality and computational cost. The findings are validated on a real-world Backblaze hard drive failure dataset, confirming the practical utility of EM acceleration. All methods are implemented in the free and open-source R package rebmix, accompanied by full source code. </dc:description><dc:date>2026</dc:date><dc:date>2026-07-22 14:36:03</dc:date><dc:type>Članek v reviji</dc:type><dc:identifier>185098</dc:identifier><dc:language>sl</dc:language></rdf:Description></rdf:RDF>
