Enhancing health survey data modeling through mixed-effects machine learning: A comparative study

Siskarossa Ika Oktora et al.

Statistical Journal of the IAOS2026https://doi.org/10.1177/18747655261418172article
ABDC C
Weight
0.50

What the paper says

Epidemiological studies frequently utilize hierarchical data structures, such as regional variations. Central obesity, commonly measured by waist circumference, is a key predictor of cardiovascular disease, one of the leading global causes of mortality. Intra-cluster correlations arising from lifestyle disparities and unequal access to healthy resources necessitate statistical models that account for both between- and within-cluster variability. Linear Mixed-Effects Models are often used for this purpose, but they may fall short in capturing nonlinearities and complex interactions among predictors. To address these limitations, tree-based extensions such as the Mixed-Effects Regression Tree and Mixed-Effects Random Forest have been introduced. MERF integrates Random Forests into the mixed-effects framework to enhance flexibility and predictive power. This study evaluates and compares the performance of LMM, MERT, and MERF in modeling waist circumference as a proxy for central obesity, using Indonesia's 2018 Basic Health Survey for West Java Province. Two modeling strategies were applied: one using full province-wide data, and another stratifying the regions into three risk categories (high, moderate, and low) based on proximity to the Special Capital Region of Jakarta. Moreover, separate models were developed for each gender to examine any difference in the contributing or influential variables for waist circumference between males and females. All models included regency/city as a random effect. Results show that MERF outperforms LMM in predictive accuracy and performs comparably to MERT, highlighting its potential for improving the modeling of clustered health survey data. Linear-based approaches, such as LMM and simple tree-based models, specifically exhibited better predictive performance than MERF for the male model. The findings support the use of mixed-effects machine learning approaches in enhancing the quality of official health statistics and informing targeted public health interventions.

Open paper page →

Cite this paper

https://doi.org/https://doi.org/10.1177/18747655261418172

Or copy a formatted citation

@article{siskarossa2026,
  title        = {{Enhancing health survey data modeling through mixed-effects machine learning: A comparative study}},
  author       = {Siskarossa Ika Oktora et al.},
  journal      = {Statistical Journal of the IAOS},
  year         = {2026},
  doi          = {https://doi.org/https://doi.org/10.1177/18747655261418172},
}

Paste directly into BibTeX, Zotero, or your reference manager.

Flag this paper

Enhancing health survey data modeling through mixed-effects machine learning: A comparative study

Flags are reviewed by the Arbiter methodology team within 5 business days.


Evidence weight

0.50

Balanced mode · F 0.40 / M 0.15 / V 0.05 / R 0.40

F · citation impact0.50 × 0.4 = 0.20
M · momentum0.50 × 0.15 = 0.07
V · venue signal0.50 × 0.05 = 0.03
R · text relevance †0.50 × 0.4 = 0.20

† Text relevance is estimated at 0.50 on the detail page — for your query’s actual relevance score, open this paper from a search result.