Why was this loan rejected? Use SHAP to explain it and LIME to cross-check
A credit scoring model works well until someone asks why one particular application was rejected. The answer then has to hold for that application.
In brief
- The CFPB requires credit denial reasons to be specific and accurate. A complex or opaque model is no exemption.
- For tree models, Tree SHAP is exact and fast, but the way it handles correlated features changes the result.
- LIME is useful as a cross-check, but it can be unstable on tabular data, and both methods can be fooled.
Picture an email from the risk team at a lender where you are deploying a model: “Application 4127 was rejected. The customer has complained. We need a specific reason by Friday.” The model is an XGBoost with a few dozen features, a good AUC and a passed backtest. Nobody in the meeting can say why this particular application was turned down.
This is a classic FDE task. The model is statistically sound, but it cannot explain itself to a single customer. In the United States, answering that question is a legal obligation for the lender, not a courtesy.
CFPB Circular 2022-03, issued on 26 May 2022, confirms that the requirement under ECOA and Regulation B to state the reasons for a denial applies to every credit decision, whatever technology is used to make it.
The CFPB goes further. Creditors may not use complex algorithms if doing so means they cannot give specific and accurate reasons. Nor can they argue that the technology they use to assess applications is too complex or too opaque to explain.
If you work for a fintech serving the US market, or for a bank that looks to this standard, the skills in this guide will be put to real use.
SHAP answers “how far did each feature push this application?”
SHAP, published by Lundberg and Lee in 2017, assigns each feature a contribution to a specific prediction. The usual way to read it is to imagine the model has a baseline, roughly the average default probability across a reference dataset.
Starting from that baseline, you look at how much each feature of the application pushes the prediction up or pulls it down. But do not assume this reading holds. Check that the baseline plus the contributions matches the probability the model returns for that application.
When the sum matches, you have an explanation the risk team can trace number by number, which is what a denial reason needs.
For tree models such as XGBoost or LightGBM, shap.TreeExplainer uses Tree SHAP, which the library’s documentation describes as a fast and exact method for trees and tree ensembles.
There is no approximation and no waiting for hours.
LIME, by Ribeiro, Singh and Guestrin in 2016, takes a different route. It generates many variants of the application by perturbing its values, asks the original model to predict each variant, and fits a simple model on this new dataset.
Samples closer to the original application get higher weight. LIME supports classifiers on tabular data, with both numeric and categorical columns, which suits credit data.
Working through application 4127
The code below computes SHAP on the probability scale in interventional mode. According to the SHAP documentation this mode needs a background dataset, so we sample 500 applications from the training set.
import shap
from lime.lime_tabular import LimeTabularExplainer
background = X_train.sample(500, random_state=0)
explainer = shap.TreeExplainer(
model,
data=background,
feature_perturbation="interventional",
model_output="probability",
)
x = X_test.loc[[4127]]
sv = explainer(x)
print(sv.base_values[0], sv.values[0].sum())
Suppose the result gives a baseline of 0.12 and a sum of contributions of +0.29, and that comparing with predict_proba confirms the model’s default probability is 0.41, above the 0.30 approval threshold the company has set.
Broken down by feature: a 52% debt-to-income ratio contributes +0.14; three late payments in the past 12 months add +0.11; only 4 months in the current job adds +0.06; income adds +0.03; and a 9-year credit history pulls it down by −0.05.
The sum: 0.14 + 0.11 + 0.06 + 0.03 − 0.05 = 0.29.
This sum is the first thing to check. If base_values + values.sum() does not match model.predict_proba, check which output you are explaining, log-odds instead of probability for example, before you read any other number.
Next, run LIME on the same application for comparison:
lime_exp = LimeTabularExplainer(
X_train.values,
feature_names=list(X_train.columns),
class_names=["tra_du", "vo_no"],
mode="classification",
random_state=0,
)
e = lime_exp.explain_instance(
x.values[0], model.predict_proba, num_features=5
)
print(e.as_list())
# [('dti > 0.45', 0.21), ('late_payments_12m > 2', 0.17),
# ('employment_months <= 6', 0.08), ('credit_age_years > 8', -0.06),
# ('income <= 9000', 0.02)]
(The class names are Vietnamese for “repaid in full” and “defaulted”.)
Compare rank and sign, not magnitude
The SHAP output and the LIME output above are measured on different scales. In this example, the SHAP values have been checked to add up exactly to the gap between the 0.12 baseline and the 0.41 prediction.
LIME weights are coefficients of a local linear model fitted on conditions such as “dti > 0.45”, so adding them up gives no meaningful number.
So when you compare, look at only two things: which features rank highest, and which direction each one pushes.
| Feature | SHAP (rank, value) | LIME (rank, weight) | Match? |
|---|---|---|---|
| Debt-to-income | 1, +0.14 | 1, +0.21 | Yes |
| Late payments, 12 months | 2, +0.11 | 2, +0.17 | Yes |
| Time in employment | 3, +0.06 | 3, +0.08 | Yes |
| Age of credit history | pulls down, −0.05 | pulls down, −0.06 | Yes |
For application 4127, the top three reasons agree on both rank and sign. You can now write to the customer: a high debt-to-income ratio, recent late payments, and a short time in the current job. Three specific statements, each tied to a number in the application.
Five steps to take at the client
First, ask the risk team which scale they want explained and what the approval threshold is, because a reason only means something next to the threshold. Then run SHAP and check that the sum matches the model’s actual output. Next, run LIME with several random seeds and compare rank and sign with SHAP.
The fourth step is the one most often skipped: test against the model itself. Lower application 4127’s debt-to-income ratio to the median and call predict_proba again.
The probability should fall clearly. If it does not, your number one reason is not as accurate as the CFPB requires. Finally, map feature names to wording customers can understand, and have the legal team approve that mapping once rather than reviewing every letter.
Three traps that produce wrong reasons that look right
The first trap is correlated features. The SHAP documentation notes that SHAP values rest on conditional expectations, so you have to choose how to handle features that depend on one another, such as income and loan amount.
interventional mode needs a background set. tree_path_dependent mode does not, because it uses the number of training samples passing through each leaf as the background distribution. The two modes can split contributions differently, so pick one and record why.
The second trap is LIME’s instability. According to Christoph Molnar’s book Interpretable Machine Learning, choosing the neighbourhood is still an unsolved problem when LIME is applied to tabular data. So do not treat a single LIME run as evidence. If the top three changes with the seed, treat it only as a hint.
The last trap is more serious. A 2019 study showed that perturbation-based explanation techniques such as LIME and SHAP can be fooled by an adversarially designed model, and the authors concluded that they are not reliable.
That is why the step above of testing with real changes to the data is mandatory, not optional.
How to show this skill on a CV
When reading fintech or bank job descriptions, look for phrases such as “model explainability”, “reason codes” and “adverse action”. On your CV, do not just write “used SHAP”.
Describe what you delivered: a pipeline that generates denial reasons for each application, with a sum check, a cross-check against LIME and counterfactual verification.
If you are applying for credit roles, be ready to walk through each step using a specific application, with the numbers you checked.
The next time an email asks why an application was rejected, there is no need to reach for slides about model accuracy. Send the sum from 0.12 to 0.41, along with evidence that every term in it has been verified.
7 sources
- Circular 2022-03: Adverse action notification requirements in connection with credit decisions based on complex algorithms (CFPB) · 2022-05-26
- "Why Should I Trust You?": Explaining the Predictions of Any Classifier · 2016-02-16
- Interpretable Machine Learning – LIME (Christoph Molnar)
- marcotcr/lime (GitHub README)
- A Unified Approach to Interpreting Model Predictions · 2017-05-22
- shap.TreeExplainer — SHAP documentation
- Fooling LIME and SHAP: Adversarial Attacks on Post hoc Explanation Methods · 2019-11-06