Confidence Intervals for the Model Performance Metrics Under the Imbalanced Classification: Evaluating SMOTE’s Impact on Metrics’ Reliability

Yury Y. Festa & Henry Penikas

Model Assisted Statistics and Applications2025https://doi.org/10.1177/15741699251350368article
ABDC C
Weight
0.41

What the paper says

AI applications in finance including those for the probability of default modeling largely involve using ML classification tools. Oversampling the very minor (very underrepresented) class of defaulted borrowers seems to be a must-be-done step always. However, by crunching more than a thousand of confidence intervals for the classification accuracy metrics, we demonstrate when such oversampling is worth engaging in. Moreover, we argue to what portion of total initial sample size such oversampling should be carried out. Our findings are valuable primarily for the credit risk modeling and Internal Ratings Based (IRB) banks, but are not limited to those and have general applications for the binary classifications in ML domain.

2 citations

Open paper page →

Cite this paper

https://doi.org/https://doi.org/10.1177/15741699251350368

Or copy a formatted citation

@article{yury2025,
  title        = {{Confidence Intervals for the Model Performance Metrics Under the Imbalanced Classification: Evaluating SMOTE’s Impact on Metrics’ Reliability}},
  author       = {Yury Y. Festa & Henry Penikas},
  journal      = {Model Assisted Statistics and Applications},
  year         = {2025},
  doi          = {https://doi.org/https://doi.org/10.1177/15741699251350368},
}

Paste directly into BibTeX, Zotero, or your reference manager.

Flag this paper

Confidence Intervals for the Model Performance Metrics Under the Imbalanced Classification: Evaluating SMOTE’s Impact on Metrics’ Reliability

Flags are reviewed by the Arbiter methodology team within 5 business days.


Evidence weight

0.41

Balanced mode · F 0.40 / M 0.15 / V 0.05 / R 0.40

F · citation impact0.25 × 0.4 = 0.10
M · momentum0.55 × 0.15 = 0.08
V · venue signal0.50 × 0.05 = 0.03
R · text relevance †0.50 × 0.4 = 0.20

† Text relevance is estimated at 0.50 on the detail page — for your query’s actual relevance score, open this paper from a search result.