Loan Approval Prediction.
Classification, tested before it is trusted
A classification study that asks more of a model than accuracy: does it generalise, where does it make mistakes, and what drives its predictions?
Predict a decision. Understand its limits.
The target is the recorded Approved or Rejected label. This study evaluates consistency with patterns in those decisions; it does not estimate repayment, default risk or borrower creditworthiness. Reliability requires explicit validation boundaries and an honest account of the errors.
Start with the applications
The supplied CSV contains 4,269 rows and 13 original columns: 2,656 Approved and 1,613 Rejected. The unique application ID and target are excluded, leaving 11 predictors: nine numeric and two categorical.
A moderately imbalanced target
Predictors include dependents, education, self-employment, annual income, requested loan amount, loan term, CIBIL score and four asset-value fields. The source report attributes the data to Kaggle but does not identify an exact version or collection provenance.
Incomplete data, explicit treatment
| Field | Missing rows | Share |
|---|---|---|
| education | 213 | 4.99% |
| self_employed | 213 | 4.99% |
| income_annum | 298 | 6.98% |
| loan_amount | 298 | 6.98% |
| cibil_score | 298 | 6.98% |
No full-row duplicates were found and all application IDs are unique. Numeric gaps use mean imputation; categorical gaps use the most frequent value. Numeric features are standardised and categorical features are one-hot encoded. Each fit learns these transformations from its own training rows.
How features differ by recorded outcome
Median observed CIBIL score is 710 for Approved and 430 for Rejected applications. Annual income and requested amount have pairwise Pearson correlation 0.93; income and luxury assets 0.93. Missing pairs are excluded. Correlated predictors complicate individual Logistic Regression coefficient interpretation and can share attribution in tree models; they are not universally irrelevant to ensembles.
Keep the validation boundary intact
- 01Raw applicationsStratified 80/20 split · seed 42
- 02Training partition3,415 rows · five outer folds
- 03Inner grid searchFive folds per candidate
- 04Fit complete pipelineImpute → encode / scale → SMOTE → classify
- 05Select model familyHighest mean outer-fold F1
The source workflow was refined so that preprocessing and resampling sit inside GridSearchCV, and model-selection folds exclude the holdout. This prevents synthetic neighbours and learned transformations from crossing an evaluation boundary. The held-out partition remains unchanged from the source split.
Ordinary SMOTE after one-hot encoding can interpolate categorical indicators. It is retained for comparison with the original study; categorical-aware sampling, class weighting and no resampling are important next comparisons.
Five models, one consistent benchmark
Held-out model comparison
| Model | Accuracy | Precision | Recall | F1 | ROC AUC | Avg precision |
|---|---|---|---|---|---|---|
| Logistic Regression | 0.9145 | 0.9473 | 0.9134 | 0.9300 | 0.9651 | 0.9805 |
| Random Forest | 0.9567 | 0.9540 | 0.9774 | 0.9656 | 0.9911 | 0.9940 |
| XGBoost | 0.9555 | 0.9642 | 0.9642 | 0.9642 | 0.9916 | 0.9947 |
| Gradient Boosting | 0.9614 | 0.9680 | 0.9699 | 0.9690 | 0.9926 | 0.9952 |
| SVC | 0.9239 | 0.9430 | 0.9341 | 0.9385 | 0.9791 | 0.9872 |
F1 balances Approved-class precision and recall under moderate imbalance. Accuracy, ROC AUC and average precision provide complementary views; none determines acceptable decision costs on its own. Average precision is computed with scikit-learn, rather than a trapezoidal PR area.
Select on validation. Report the test.
Training-only nested validation
Random Forest has the highest mean outer-fold F1 in this run. The family-selection record was saved before holdout evaluation. A different family can lead an individual holdout metric; that does not retrospectively change the selection. Fold variability describes sensitivity within this dataset; it is not a confidence interval or proof of performance on a new lending population.
Default versus tuned models
For the selected Random Forest, tuning changes mean outer-fold F1 by +0.0003. The benchmark provides stronger evidence for choosing a model family than for an extensive search within that family; these differences are not formal significance tests.
| Model | Default F1 | Tuned F1 | Change |
|---|---|---|---|
| Logistic Regression | 0.9090 | 0.9077 | -0.0014 |
| Random Forest | 0.9538 | 0.9540 | +0.0003 |
| XGBoost | 0.9486 | 0.9495 | +0.0009 |
| Gradient Boosting | 0.9491 | 0.9523 | +0.0032 |
| SVC | 0.9219 | 0.9249 | +0.0030 |
Where the selected model gets it wrong
Selected-model error analysis
Approved-class precision is 0.9540 and recall is 0.9774. A false approval means a recorded Rejected application was predicted Approved; a false rejection reverses that mismatch. Without repayment outcomes, these errors cannot be equated with risky borrowers, viable borrowers or monetary losses.
How well do the models rank outcomes?
What is the precision–recall trade-off?
Explain the model, not the lending policy
What drives the selected model?
The plot measures average prediction contributions, not causal effects or the merit of an application. Correlated income, loan and asset fields can share predictive information; importance rankings depend on the model and explanation assumptions. Numeric reconstruction of predicted probabilities was checked for additivity.
Strong classification is not a lending mandate
- Historical labels may encode previous policy. Approval decisions are not repayment outcomes or proof of creditworthiness.
- Fairness has not been assessed. Omitted protected characteristics do not remove proxy bias, and SHAP does not establish equitable outcomes.
- Collection provenance, exact source version and decision policy are not documented in the supplied materials. Missingness mechanisms are unknown.
- The source notebook previously explored the complete dataset. The revised holdout is excluded from fitting and selection, but is not a new prospective sample.
- There is no external or temporal validation. Population shift, calibration and threshold suitability remain untested.
- Any real lending application would require an appropriate fairness assessment, policy and regulatory review, monitoring and human oversight.
A result that can be reproduced
Training, preprocessing, evaluation, plotting and explanation are separated into small Python modules. The selected artefact contains the complete preprocessing and classification pipeline. Seeded splits, exact grids, aggregate fold results, dependency versions and dataset checksum accompany the figures.
Reproduce the evidence
python -m pip install -r requirements.txt
python -m src.train
python -m src.figures
python -m src.explain
python -m src.verifyThe next experiment should challenge the result
- Compare SMOTENC, class weights and no sampling under identical validation splits.
- Evaluate probability calibration and thresholds against explicitly defined error costs.
- Test temporal or external generalisation using an independently documented dataset.
- Assess group outcomes and proxy sensitivity when suitable attributes and governance are available.
- Inspect representative errors and local SHAP explanations alongside the aggregate benchmark.