# How to pick a classification threshold by the cost of each error, not the 0.5 default

> In a credit-scoring example from scikit-learn, the model stays the same and only the decision threshold moves. That one change makes the business outcome nearly twice as good.

Bản gốc: https://fdetimes.net/en/guides/choose-classification-threshold-by-error-cost/

In one credit-scoring example, the scikit-learn documentation keeps the model fixed and changes a single number: the decision threshold. The business gain moves from -209 to -143, which is nearly twice as good. There are no new features and no retraining.

If you want to work as an FDE, practise this before you meet clients. Data scientists tend to report AUC or accuracy. Clients want to know how accurate is accurate enough, and how much money or how many staff hours each mistake costs them. This guide walks through each step so you can answer that question in code, on your own laptop.

**What you will do:** write a cost function, sweep thresholds on a validation set, pick the best threshold and check it on a test set. **What you need:** Python 3 and any binary classifier that outputs probability scores. The code below uses plain Python with no libraries, so you can see every calculation. It is simplified for learning.

## Are you choosing a metric, or choosing which error to avoid?

Two definitions to keep in mind. Precision is the share of positive predictions that are actually positive. Recall is the share of real positives the model catches. Domino's data science dictionary makes a point that is easy to miss: deciding to prioritise precision or recall changes which type of error the model is optimised to avoid.

Google's ML Crash Course gives a practical rule. Prioritise recall when false negatives cost more than false positives. Prioritise precision when positive predictions must be correct. Do not use accuracy on imbalanced data.

Why not accuracy? Picture 1,000 loan applications, 100 of them from bad borrowers. A "lazy" model that labels everyone a good borrower scores 90% accuracy and catches no bad borrowers at all. So the first step is a conversation with the client. Code comes later.

## Step 1: get the cost of each error from the client

In the scikit-learn example, "positive" means a bad borrower. Lending to a bad borrower by mistake (a false negative) costs on average five times as much as wrongly rejecting a good one (a false positive). The gain matrix therefore gives -1 for each FP and -5 for each FN.

```python
COST_FP = 1   # good borrower wrongly rejected
COST_FN = 5   # bad borrower wrongly approved

def business_cost(tp, fp, tn, fn):
return COST_FP * fp + COST_FN * fn
```

**Check:** does the client agree with the 1:5 ratio? The number does not have to be exact, but it has to come from the people who know the business. Engineers should not guess it.

## Step 2: turn scores into labels

The model returns probabilities. To get labels you need a cut-off probability, called the classification threshold. Google notes that different thresholds usually produce different numbers of TP, FP, TN and FN.

```python
def confusion_at(y_true, y_score, threshold):
tp = fp = tn = fn = 0
for y, s in zip(y_true, y_score):
pred = 1 if s >= threshold else 0
if pred == 1 and y == 1: tp += 1
elif pred == 1 and y == 0: fp += 1
elif pred == 0 and y == 0: tn += 1
else: fn += 1
return tp, fp, tn, fn
```

**Check:** tp + fp + tn + fn must equal the number of samples. If it doesn't, your labels contain values other than 0 and 1.

## Step 3: sweep thresholds instead of trusting 0.5

The scikit-learn documentation says plainly that the default strategy (a 0.5 threshold) is probably not optimal for the problem at hand. The best threshold is the one that gives the best value of the metric you chose. Here, that metric is business cost.

```python
def sweep(y_true, y_score):
rows = []
for i in range(5, 96, 5):
t = i / 100
tp, fp, tn, fn = confusion_at(y_true, y_score, t)
precision = tp / (tp + fp) if tp + fp else 0.0
recall = tp / (tp + fn) if tp + fn else 0.0
rows.append((t, business_cost(tp, fp, tn, fn), precision, recall))
return sorted(rows, key=lambda r: r[1])
```

To see what the output looks like, go back to the 1,000 hypothetical applications above. The figures in the table are illustrative, not real measurements.

| Classification | TP / FP / FN | Precision | Recall | Accuracy | Cost |
|---|---|---|---|---|---|
| Label everyone a good borrower | 0 / 0 / 100 | — | 0% | 90% | 500 |
| Threshold 0.5 | 40 / 10 / 60 | 80% | 40% | 93% | 310 |
| Threshold 0.2 | 75 / 60 / 25 | 55.6% | 75% | 91.5% | 185 |

Look at what happens between 0.5 and 0.2: accuracy falls, but cost falls sharply. Recall goes up and precision goes down. scikit-learn reports the same kind of trade-off when it tunes the threshold in its credit example. If you report accuracy, you will pick the wrong threshold.

## Step 4: never tune the threshold on training data

The scikit-learn guide says you should never use the same data to train the classifier and to tune the threshold, because it leads to overfitting. On the training set the model is more confident than it really is, so any threshold found there will be too optimistic.

The safe approach is a three-way split: train to fit the model, validation to run `sweep`, and test to measure the cost one last time at the chosen threshold. If the test cost is clearly higher than the validation cost, your threshold is fitting noise.

**Check:** run `confusion_at` on the test set with the chosen threshold, then compute `business_cost`. Put this number in your report, not the validation number.

## So what is AUC for?

AUC and the ROC curve show how well the model separates the two classes across every possible threshold. That is exactly why AUC cannot tell you where to set the threshold. It is useful for comparing two models before you choose a threshold, but it does not replace the cost figure at the threshold you will actually run.

**Điểm mấu chốt:** AUC helps you choose a model. Cost at a specific threshold helps you make a decision. Clients pay for the second.

## What this looks like on a client site

Imagine you have deployed a model that flags suspicious transactions for a bank. The operations team complains about too many false alarms. The risk team worries about missed cases. Both are right, and the argument only stops when everyone works from one shared figure: the cost of each type of error.

Start by sitting down with both teams and asking: how many times more does one missed case cost than one false alarm? Then run `sweep`, show them a table like the one above and let them pick the row that suits them. The threshold becomes a business decision with evidence behind it, rather than a technical parameter hidden in the code.

On your CV and in interviews, don't just write "achieved 0.9 AUC". Write that you turned the client's requirements into a cost function, chose a threshold on a separate validation set, and cut the cost of errors by a given percentage compared with the default threshold.

When a job description asks for someone who can translate business needs into metrics, this example is evidence that you can.

The best model in your notebook can still be the most expensive one in production if the threshold is wrong. The engineers who get invited back are often the ones who ask the client "which mistake costs more?" before they run any code.

**Thử ngay tuần này:**

- Take a classifier you already have, run the sweep function from this guide with an FP:FN cost ratio of 1:5, and compare the best threshold with 0.5.
- Prepare a single question for your next client meeting: 'How many times more does missing one bad case cost than raising one false alarm on a good case?'
- Add a table to your project README showing precision, recall and cost at the chosen threshold, with one line explaining why you chose it.

## Nguồn

- [What is Model Evaluation (Domino Data Science Dictionary)](https://domino.ai/data-science-dictionary/model-evaluation)

- [Thresholds and the confusion matrix (Google ML Crash Course)](https://developers.google.com/machine-learning/crash-course/classification/thresholding)

- [Accuracy, precision, and recall (Google ML Crash Course)](https://developers.google.com/machine-learning/crash-course/classification/accuracy-precision-recall)

- [Post-tuning the decision threshold for cost-sensitive learning (scikit-learn)](https://scikit-learn.org/stable/auto_examples/model_selection/plot_cost_sensitive_learning.html)

- [Tuning the decision threshold for class prediction (scikit-learn User Guide)](https://scikit-learn.org/stable/modules/classification_threshold.html)
