Generating realistic synthetic patient cohorts: enforcing statistical distributions, correlations, and logical constraints

Fasseeh, Ahmad Nader; Ashmawy, Rasha; Hren, Rok; ElFass, Kareem; Imre, Attila; Németh, Bertalan; Nagy, Dávid; Nagy, Balázs; Vokó, Zoltán

Repository of the University of Ljubljana

Details

Generating realistic synthetic patient cohorts: enforcing statistical distributions, correlations, and logical constraints
ID Fasseeh, Ahmad Nader (Author), ID Ashmawy, Rasha (Author), ID Hren, Rok (Author), ID ElFass, Kareem (Author), ID Imre, Attila (Author), ID Németh, Bertalan (Author), ID Nagy, Dávid (Author), ID Nagy, Balázs (Author), ID Vokó, Zoltán (Author)

	PDF - Presentation file, Download (1,11 MB) MD5: E072E3AA53CE04C567D7D046712BF8F6
	URL - Source URL, Visit https://www.mdpi.com/1999-4893/18/8/475

Abstract

Large, high-quality patient datasets are essential for applications like economic modeling and patient simulation. However, real-world data is often inaccessible or incomplete. Synthetic patient data offers an alternative, and current methods often fail to preserve clinical plausibility, real-world correlations, and logical consistency. This study presents a patient cohort generator designed to produce realistic, statistically valid synthetic datasets. The generator uses predefined probability distributions and Cholesky decomposition to reflect real-world correlations. A dependency matrix handles variable relationships in the right order. Hard limits block unrealistic values, and binary variables are set using percentiles to match expected rates. Validation used two datasets, NHANES (2021–2023) and the Framingham Heart Study, evaluating cohort diversity (general, cardiac, low-dimensional), data sparsity (five correlation scenarios), and model performance (MSE, RMSE, R2, SSE, correlation plots). Results demonstrated strong alignment with real-world data in central tendency, dispersion, and correlation structures. Scenario A (empirical correlations) performed best (R2 = 86.8–99.6%, lowest SSE and MAE). Scenario B (physician-estimated correlations) also performed well, especially in a low-dimensions population (R2 = 80.7%). Scenario E (no correlation) performed worst. Overall, the proposed model provides a scalable, customizable solution for generating synthetic patient cohorts, supporting reliable simulations and research when real-world data is limited. While deep learning approaches have been proposed for this task, they require access to large-scale real datasets and offer limited control over statistical dependencies or clinical logic. Our approach addresses this gap.

Language:	English
Keywords:	healthcare, health economics, health informatics, synthetic data
Work type:	Article
Typology:	1.01 - Original Scientific Article
Organization:	FMF - Faculty of Mathematics and Physics
Publication status:	Published
Publication version:	Version of Record
Year:	2025
Number of pages:	29 str.
Numbering:	Vol. 18, iss. 8, art. no. 475
PID:	20.500.12556/RUL-171095
UDC:	614
ISSN on article:	1999-4893
DOI:	10.3390/a18080475
COBISS.SI-ID:	244713731
Publication date in RUL:	04.08.2025
Views:	246
Downloads:	50
Metadata:
:	Copy citation
Share:

Record is a part of a journal

Title:	Algorithms
Shortened title:	Algorithms
Publisher:	MDPI
ISSN:	1999-4893
COBISS.SI-ID:	517501977

Secondary language

Language:	Slovenian
Keywords:	zdravstvo, zdravstvena ekonomija, zdravstvena informatika, sintetični podatki

Similar works from RUL:
Similar works from other Slovenian collections:

Details

Record is a part of a journal

Secondary language

Similar documents