← All projectsPROJECT 004 / ANALYSED

Loan Approval Prediction.

Classification, tested before it is trusted

A classification study that asks more of a model than accuracy: does it generalise, where does it make mistakes, and what drives its predictions?

APPLICATIONS4,269
MODELS5
SELECTED MODELRandom Forest
CV F10.9540Five outer folds · SD 0.0038
TEST F10.9656
TEST ROWS854
01 / CASE STUDY

Predict a decision. Understand its limits.

The target is the recorded Approved or Rejected label. This study evaluates consistency with patterns in those decisions; it does not estimate repayment, default risk or borrower creditworthiness. Reliability requires explicit validation boundaries and an honest account of the errors.

02 / CASE STUDY

Start with the applications

The supplied CSV contains 4,269 rows and 13 original columns: 2,656 Approved and 1,613 Rejected. The unique application ID and target are excluded, leaving 11 predictors: nine numeric and two categorical.

A moderately imbalanced target

A moderately imbalanced target
A moderately imbalanced target. Extracted from the source notebook: 2,656 Approved and 1,613 Rejected applications. Approved is the positive class throughout.
Extracted from the source notebook: 2,656 Approved and 1,613 Rejected applications. Approved is the positive class throughout.

Predictors include dependents, education, self-employment, annual income, requested loan amount, loan term, CIBIL score and four asset-value fields. The source report attributes the data to Kaggle but does not identify an exact version or collection provenance.

03 / CASE STUDY

Incomplete data, explicit treatment

Missingness verified directly from the supplied CSV
FieldMissing rowsShare
education2134.99%
self_employed2134.99%
income_annum2986.98%
loan_amount2986.98%
cibil_score2986.98%

No full-row duplicates were found and all application IDs are unique. Numeric gaps use mean imputation; categorical gaps use the most frequent value. Numeric features are standardised and categorical features are one-hot encoded. Each fit learns these transformations from its own training rows.

How features differ by recorded outcome

How features differ by recorded outcome
How features differ by recorded outcome. Extracted from the original notebook: distributions of CIBIL score, annual income, requested loan amount and term by recorded decision. These full-dataset plots are descriptive, rather than independent model-validation evidence.
Extracted from the original notebook: distributions of CIBIL score, annual income, requested loan amount and term by recorded decision. These full-dataset plots are descriptive, rather than independent model-validation evidence.

Median observed CIBIL score is 710 for Approved and 430 for Rejected applications. Annual income and requested amount have pairwise Pearson correlation 0.93; income and luxury assets 0.93. Missing pairs are excluded. Correlated predictors complicate individual Logistic Regression coefficient interpretation and can share attribution in tree models; they are not universally irrelevant to ensembles.

04 / CASE STUDY

Keep the validation boundary intact

Leakage-contained classification pipelineVERIFIED
  1. 01Raw applicationsStratified 80/20 split · seed 42
  2. 02Training partition3,415 rows · five outer folds
  3. 03Inner grid searchFive folds per candidate
  4. 04Fit complete pipelineImpute → encode / scale → SMOTE → classify
  5. 05Select model familyHighest mean outer-fold F1
Untouched holdout854 original rows · final evaluation
Saved pipelinePreprocessing + sampler + classifier
SHAPInterpretation after selection
SMOTE only resamples fitting data. Inner validation, outer validation and holdout rows remain original observations.

The source workflow was refined so that preprocessing and resampling sit inside GridSearchCV, and model-selection folds exclude the holdout. This prevents synthetic neighbours and learned transformations from crossing an evaluation boundary. The held-out partition remains unchanged from the source split.

Ordinary SMOTE after one-hot encoding can interpolate categorical indicators. It is retained for comparison with the original study; categorical-aware sampling, class weighting and no resampling are important next comparisons.

05 / CASE STUDY

Five models, one consistent benchmark

Held-out model comparison

Held-out model comparison
Held-out model comparison. All five models use the same 854 original holdout rows. Random Forest was selected before these results were evaluated.
All five models use the same 854 original holdout rows. Random Forest was selected before these results were evaluated.
Held-out results · Approved = positive class
ModelAccuracyPrecisionRecallF1ROC AUCAvg precision
Logistic Regression0.91450.94730.91340.93000.96510.9805
Random Forest0.95670.95400.97740.96560.99110.9940
XGBoost0.95550.96420.96420.96420.99160.9947
Gradient Boosting0.96140.96800.96990.96900.99260.9952
SVC0.92390.94300.93410.93850.97910.9872

F1 balances Approved-class precision and recall under moderate imbalance. Accuracy, ROC AUC and average precision provide complementary views; none determines acceptable decision costs on its own. Average precision is computed with scikit-learn, rather than a trapezoidal PR area.

06 / CASE STUDY

Select on validation. Report the test.

Training-only nested validation

Training-only nested validation
Training-only nested validation. Random Forest: mean F1 0.9540 ± 0.0038 population SD across five outer folds. Each outer fold contains a separate five-fold tuning search.
Random Forest: mean F1 0.9540 ± 0.0038 population SD across five outer folds. Each outer fold contains a separate five-fold tuning search.

Random Forest has the highest mean outer-fold F1 in this run. The family-selection record was saved before holdout evaluation. A different family can lead an individual holdout metric; that does not retrospectively change the selection. Fold variability describes sensitivity within this dataset; it is not a confidence interval or proof of performance on a new lending population.

Default versus tuned models

Default versus tuned models
Default versus tuned models. Mean F1 compares default and tuned classifiers on the same five outer folds. The size and direction of tuning gains should be read from the results, rather than assumed.
Mean F1 compares default and tuned classifiers on the same five outer folds. The size and direction of tuning gains should be read from the results, rather than assumed.

For the selected Random Forest, tuning changes mean outer-fold F1 by +0.0003. The benchmark provides stronger evidence for choosing a model family than for an extensive search within that family; these differences are not formal significance tests.

Tuning effect on training-only outer validation
ModelDefault F1Tuned F1Change
Logistic Regression0.90900.9077-0.0014
Random Forest0.95380.9540+0.0003
XGBoost0.94860.9495+0.0009
Gradient Boosting0.94910.9523+0.0032
SVC0.92190.9249+0.0030
07 / CASE STUDY

Where the selected model gets it wrong

Selected-model error analysis

Selected-model error analysis
Selected-model error analysis. 519 true approvals and 298 true rejections; 25 false approvals and 12 false rejections on 854 applications.
519 true approvals and 298 true rejections; 25 false approvals and 12 false rejections on 854 applications.

Approved-class precision is 0.9540 and recall is 0.9774. A false approval means a recorded Rejected application was predicted Approved; a false rejection reverses that mismatch. Without repayment outcomes, these errors cannot be equated with risky borrowers, viable borrowers or monetary losses.

How well do the models rank outcomes?

How well do the models rank outcomes?
How well do the models rank outcomes?. ROC AUC for the selected model is 0.9911. Curves show true-positive versus false-positive rates across thresholds.
ROC AUC for the selected model is 0.9911. Curves show true-positive versus false-positive rates across thresholds.

What is the precision–recall trade-off?

What is the precision–recall trade-off?
What is the precision–recall trade-off?. The selected model has average precision 0.9940. The reference line is the observed approval prevalence in the holdout; changing a threshold trades precision against recall.
The selected model has average precision 0.9940. The reference line is the observed approval prevalence in the holdout; changing a threshold trades precision against recall.
08 / CASE STUDY

Explain the model, not the lending policy

What drives the selected model?

What drives the selected model?
What drives the selected model?. Top transformed features by mean absolute SHAP: cibil score, loan term, loan amount. Tree-path-dependent Tree SHAP explains Approved-class probability across all 854 holdout applications. Background: Fitted tree leaf frequencies, including SMOTE training observations.
Top transformed features by mean absolute SHAP: cibil score, loan term, loan amount. Tree-path-dependent Tree SHAP explains Approved-class probability across all 854 holdout applications. Background: Fitted tree leaf frequencies, including SMOTE training observations.

The plot measures average prediction contributions, not causal effects or the merit of an application. Correlated income, loan and asset fields can share predictive information; importance rankings depend on the model and explanation assumptions. Numeric reconstruction of predicted probabilities was checked for additivity.

09 / CASE STUDY

Strong classification is not a lending mandate

  • Historical labels may encode previous policy. Approval decisions are not repayment outcomes or proof of creditworthiness.
  • Fairness has not been assessed. Omitted protected characteristics do not remove proxy bias, and SHAP does not establish equitable outcomes.
  • Collection provenance, exact source version and decision policy are not documented in the supplied materials. Missingness mechanisms are unknown.
  • The source notebook previously explored the complete dataset. The revised holdout is excluded from fitting and selection, but is not a new prospective sample.
  • There is no external or temporal validation. Population shift, calibration and threshold suitability remain untested.
  • Any real lending application would require an appropriate fairness assessment, policy and regulatory review, monitoring and human oversight.
10 / CASE STUDY

A result that can be reproduced

Training, preprocessing, evaluation, plotting and explanation are separated into small Python modules. The selected artefact contains the complete preprocessing and classification pipeline. Seeded splits, exact grids, aggregate fold results, dependency versions and dataset checksum accompany the figures.

Reproduce the evidence

python -m pip install -r requirements.txt
python -m src.train
python -m src.figures
python -m src.explain
python -m src.verify
Python 3.11 · supplied CSV required. The source download excludes the virtual environment, administrative report and redundant model binaries.
11 / CASE STUDY

The next experiment should challenge the result

  • Compare SMOTENC, class weights and no sampling under identical validation splits.
  • Evaluate probability calibration and thresholds against explicitly defined error costs.
  • Test temporal or external generalisation using an independently documented dataset.
  • Assess group outcomes and proxy sensitivity when suitable attributes and governance are available.
  • Inspect representative errors and local SHAP explanations alongside the aggregate benchmark.
04 / LET’S TALK

Got a difficult
dataset?

Let’s figure out what
it is trying to tell us.