Influence of sampling methods on bankruptcy prediction: normal vs. abnormal economic conditions
A.M.Z Huq & Wonder Mahembe
What the paper says
Bankruptcy prediction research has largely emphasised model performance through feature selection and algorithm optimisation, while the equally important challenge of class imbalance remains underexplored. Most studies also focus on publicly listed firms, reflecting the accessibility of standardised data. Our study makes a novel and valuable contribution by leveraging a large-scale dataset of private firms - an economically significant yet understudied segment. Using 2,039,222 firm-year observations from 430,800 private firms between 2012 and 2021, we evaluate four machine learning models, five sampling techniques, and two distinct economic periods. Results show that sampling choice strongly influences accuracy and feature relevance, depending on macroeconomic conditions. Importantly, simple interpretable models built on theoretically grounded features (e.g., Altman, 1968) achieve robust predictions, challenging prevailing reliance on complex methods, while Extreme Gradient Boosting (XGBoost) consistently outperforms alternatives. By focusing on private firms, the study provides unique insights and underscores methodological choices crucial for reliable bankruptcy prediction.
Evidence weight
Balanced mode · F 0.40 / M 0.15 / V 0.05 / R 0.40
| F · citation impact | 0.50 × 0.4 = 0.20 |
| M · momentum | 0.50 × 0.15 = 0.07 |
| V · venue signal | 0.50 × 0.05 = 0.03 |
| R · text relevance † | 0.50 × 0.4 = 0.20 |
† Text relevance is estimated at 0.50 on the detail page — for your query’s actual relevance score, open this paper from a search result.