FDE PulseFDE jobs open 434New in the last 7 days 27
VI

The newspaper of the Forward Deployed Engineer

Guides

How to pick a classification threshold by the cost of each error, not the 0.5 default

In a credit-scoring example from scikit-learn, the model stays the same and only the decision threshold moves. That one change makes the business outcome nearly twice as good.

Ảnh cận cảnh laptop hiển thị mã Python trên bàn làm việc, gợi không khí kỹ sư đang thử nghiệm model trước khi gặp khách hàng.
Photo: Negative Space / CC0

In brief

  • Choosing between precision and recall really means choosing which error the model is optimised to avoid. The answer depends on what each error costs the client.
  • The default 0.5 threshold is probably not the best one for a real problem. In the scikit-learn example, tuning the threshold to cost moved the business gain from -209 to -143.
  • AUC is for comparing models because it covers every threshold. A decision needs one threshold, and you should tune it on data that was not used for training.
ShareLinkedInFacebookX

In one credit-scoring example, the scikit-learn documentation keeps the model fixed and changes a single number: the decision threshold. The business gain moves from -209 to -143, which is nearly twice as good. There are no new features and no retraining.

If you want to work as an FDE, practise this before you meet clients. Data scientists tend to report AUC or accuracy. Clients want to know how accurate is accurate enough, and how much money or how many staff hours each mistake costs them. This guide walks through each step so you can answer that question in code, on your own laptop.

What you will do: write a cost function, sweep thresholds on a validation set, pick the best threshold and check it on a test set. What you need: Python 3 and any binary classifier that outputs probability scores. The code below uses plain Python with no libraries, so you can see every calculation. It is simplified for learning.

Are you choosing a metric, or choosing which error to avoid?

Two definitions to keep in mind. Precision is the share of positive predictions that are actually positive. Recall is the share of real positives the model catches. Domino’s data science dictionary makes a point that is easy to miss: deciding to prioritise precision or recall changes which type of error the model is optimised to avoid.

Google’s ML Crash Course gives a practical rule. Prioritise recall when false negatives cost more than false positives. Prioritise precision when positive predictions must be correct. Do not use accuracy on imbalanced data.

Why not accuracy? Picture 1,000 loan applications, 100 of them from bad borrowers. A “lazy” model that labels everyone a good borrower scores 90% accuracy and catches no bad borrowers at all. So the first step is a conversation with the client. Code comes later.

Step 1: get the cost of each error from the client

In the scikit-learn example, “positive” means a bad borrower. Lending to a bad borrower by mistake (a false negative) costs on average five times as much as wrongly rejecting a good one (a false positive). The gain matrix therefore gives -1 for each FP and -5 for each FN.

COST_FP = 1   # good borrower wrongly rejected
COST_FN = 5   # bad borrower wrongly approved

def business_cost(tp, fp, tn, fn):
    return COST_FP * fp + COST_FN * fn

Check: does the client agree with the 1:5 ratio? The number does not have to be exact, but it has to come from the people who know the business. Engineers should not guess it.

Step 2: turn scores into labels

The model returns probabilities. To get labels you need a cut-off probability, called the classification threshold. Google notes that different thresholds usually produce different numbers of TP, FP, TN and FN.

def confusion_at(y_true, y_score, threshold):
    tp = fp = tn = fn = 0
    for y, s in zip(y_true, y_score):
        pred = 1 if s >= threshold else 0
        if pred == 1 and y == 1: tp += 1
        elif pred == 1 and y == 0: fp += 1
        elif pred == 0 and y == 0: tn += 1
        else: fn += 1
    return tp, fp, tn, fn

Check: tp + fp + tn + fn must equal the number of samples. If it doesn’t, your labels contain values other than 0 and 1.

Step 3: sweep thresholds instead of trusting 0.5

The scikit-learn documentation says plainly that the default strategy (a 0.5 threshold) is probably not optimal for the problem at hand. The best threshold is the one that gives the best value of the metric you chose. Here, that metric is business cost.

def sweep(y_true, y_score):
    rows = []
    for i in range(5, 96, 5):
        t = i / 100
        tp, fp, tn, fn = confusion_at(y_true, y_score, t)
        precision = tp / (tp + fp) if tp + fp else 0.0
        recall = tp / (tp + fn) if tp + fn else 0.0
        rows.append((t, business_cost(tp, fp, tn, fn), precision, recall))
    return sorted(rows, key=lambda r: r[1])

To see what the output looks like, go back to the 1,000 hypothetical applications above. The figures in the table are illustrative, not real measurements.

Classification TP / FP / FN Precision Recall Accuracy Cost
Label everyone a good borrower 0 / 0 / 100 — 0% 90% 500
Threshold 0.5 40 / 10 / 60 80% 40% 93% 310
Threshold 0.2 75 / 60 / 25 55.6% 75% 91.5% 185

Look at what happens between 0.5 and 0.2: accuracy falls, but cost falls sharply. Recall goes up and precision goes down. scikit-learn reports the same kind of trade-off when it tunes the threshold in its credit example. If you report accuracy, you will pick the wrong threshold.

Step 4: never tune the threshold on training data

The scikit-learn guide says you should never use the same data to train the classifier and to tune the threshold, because it leads to overfitting. On the training set the model is more confident than it really is, so any threshold found there will be too optimistic.

The safe approach is a three-way split: train to fit the model, validation to run sweep, and test to measure the cost one last time at the chosen threshold. If the test cost is clearly higher than the validation cost, your threshold is fitting noise.

Check: run confusion_at on the test set with the chosen threshold, then compute business_cost. Put this number in your report, not the validation number.

So what is AUC for?

AUC and the ROC curve show how well the model separates the two classes across every possible threshold. That is exactly why AUC cannot tell you where to set the threshold. It is useful for comparing two models before you choose a threshold, but it does not replace the cost figure at the threshold you will actually run.

What this looks like on a client site

Imagine you have deployed a model that flags suspicious transactions for a bank. The operations team complains about too many false alarms. The risk team worries about missed cases. Both are right, and the argument only stops when everyone works from one shared figure: the cost of each type of error.

Start by sitting down with both teams and asking: how many times more does one missed case cost than one false alarm? Then run sweep, show them a table like the one above and let them pick the row that suits them. The threshold becomes a business decision with evidence behind it, rather than a technical parameter hidden in the code.

On your CV and in interviews, don’t just write “achieved 0.9 AUC”. Write that you turned the client’s requirements into a cost function, chose a threshold on a separate validation set, and cut the cost of errors by a given percentage compared with the default threshold.

When a job description asks for someone who can translate business needs into metrics, this example is evidence that you can.

The best model in your notebook can still be the most expensive one in production if the threshold is wrong. The engineers who get invited back are often the ones who ask the client “which mistake costs more?” before they run any code.

5 sources
Read next on the roadmap · Stage 6: MeasurementLLM-as-judge in practice: calibrating a judge against the client's expert pass/fail labelsA judge with 90% agreement can still miss every wrong answer. If you want a client to trust your eval numbers, measure the judge against the person on their side who knows the domain best.