# A one-line prompt edit is a deploy: running a 5-20-50-100% canary at the client

> A prompt edit that only changes formatting can shift accuracy by around 5%. At a client, every change like that needs a feature flag to split traffic, a control group to compare against and a kill switch that is always ready.

Original: https://fdetimes.net/en/guides/canary-rollout-llm-prompt-changes/

Picture this: a client asks you to tweak the prompt so their support chatbot sounds "more polite". You add a few greetings, switch bullet points to a numbered list and test it in staging. Everything looks fine. Two days later, the customer service team reports that the bot is getting the refund policy wrong.

This scenario is hypothetical, but not far-fetched. LaunchDarkly warns that a small prompt edit or a model swap can work well in staging and still degrade the user experience in production. An analysis on the tianpan.co blog found that minor changes to prompt formatting shifted accuracy by around 5%.

So switching bullets to numbers is a deploy. This guide walks you through building the whole process yourself: a feature flag that splits traffic between the old and new prompts, logging to compare the two groups, a rollback decision function and a schedule that ramps up through 5-20-50-100%.

The code here is illustrative Python, trimmed down and not tied to any vendor's SDK. In a real engagement, swap it for whatever flagging tool the client already uses.

## A canary is really a comparison

The Google SRE Workbook defines canarying as deploying a change to a portion of a service for a limited time and evaluating that change. The portion that receives the change is the canary; the rest is the control. Without a control group alongside it, you are only running a trial, not a canary.

**What you need:** Python 3, an existing function that calls your LLM (called `call_llm` here), a queryable place to write logs, and two clearly named prompt versions, for example `v1.3` and `v2.0`.

## Step 1: Bundle prompt and model into variations

Do not edit the prompt string directly in code. Declare each configuration as a named variation, so that rolling back means changing a selection rather than redeploying.

```python
FLAG = {
    "name": "support_bot_prompt_v2",
    "enabled": True,           # emergency kill switch
    "rollout_percent": 5,
    "internal_users": {"qa@client.vn"},
}

VARIATIONS = {
    "control": {"model": "current-model", "prompt_version": "v1.3"},
    "canary":  {"model": "current-model", "prompt_version": "v2.0"},
}
```

(The comment on `enabled` marks it as the emergency off switch; `current-model` means "current model".) In this example the two variations differ only in the prompt. If you want to change the model, do it under a separate flag. When prompt and model change together, you cannot tell which one caused the result.

## Step 2: Split users stably, internal users first

Flagsmith describes its own process for LLM features like this: first enable the feature in production for internal users only, then open it to 5% of users, then 50% to run an A/B test, and only then go to 100%. The code below follows that order.

```python
import hashlib

def bucket(user_id: str, flag_name: str) -> int:
    h = hashlib.sha256(f"{flag_name}:{user_id}".encode()).hexdigest()
    return int(h[:8], 16) % 100

def pick_variation(user_id: str, email: str) -> str:
    if not FLAG["enabled"]:
        return "control"
    if email in FLAG["internal_users"]:
        return "canary"
    if bucket(user_id, FLAG["name"]) < FLAG["rollout_percent"]:
        return "canary"
    return "control"
```

**Check:** run the function on 10,000 fake user_ids and count how many land in the canary; the result should be around 500. Calling it again with the same user_id must always return the same variation. Without that stability, one person could get an answer from the old prompt on one turn and the new prompt on the next, and your data becomes noisy.

## Step 3: Log at the right moment

FeatBit recommends recording the variation at the moment the LLM actually runs, not when the page loads. It also recommends logging `promptVersion` alongside it, so that if someone quietly edits the prompt mid-rollout, the change does not contaminate the canary results.

```python
import time

def answer(user_id, email, question):
    variation = pick_variation(user_id, email)
    cfg = VARIATIONS[variation]
    started = time.time()
    reply = call_llm(cfg["model"], load_prompt(cfg["prompt_version"]), question)
    log_event({
        "user_id": user_id,
        "flag": FLAG["name"],
        "variation": variation,
        "prompt_version": cfg["prompt_version"],
        "model": cfg["model"],
        "latency_ms": int((time.time() - started) * 1000),
    })
    return reply
```

`call_llm`, `load_prompt` and `log_event` are functions you already have. **Check:** every log line must contain all four fields: variation, prompt_version, model and latency. Missing even one means you cannot compare the two groups reliably. In practice, also log the question type (for example `refund`) so the next step can count samples per type.

**Key point:** A one-line prompt edit is a deploy: it needs a flag, a control group and a kill switch at the ready.

## Step 4: Track a few metrics, set thresholds in advance

Google SRE advises choosing only the few most important metrics to evaluate a canary, probably no more than a dozen. For a customer support chatbot, you might start with error rate, latency, the rate at which users have to be handed off to a human agent, and a quality score graded on a sample of conversations.

The function below returns one of three decisions: roll back, wait, or promote. Guardrails are checked first, because a clear failure should stop things immediately. To promote, the canary must also have collected enough samples, both in total and for rare question types.

```python
MIN_TOTAL = 500    # minimum total conversations in the canary
MIN_REFUND = 20    # minimum number of refund questions

def decide(canary: dict, control: dict) -> str:
    if canary["error_rate"] > control["error_rate"] * 1.5:
        return "rollback"
    if canary["p95_latency_ms"] > control["p95_latency_ms"] * 1.3:
        return "rollback"
    if canary["handoff_rate"] > control["handoff_rate"] + 0.03:
        return "rollback"
    if canary["quality_score"] < control["quality_score"] - 0.05:
        return "rollback"
    if canary["n_total"] < MIN_TOTAL or canary["n_refund"] < MIN_REFUND:
        return "wait"
    return "promote"
```

(`MIN_TOTAL` is the minimum total number of canary conversations; `MIN_REFUND` the minimum number of refund questions.) Every threshold above is only an example. Agree the real thresholds with the client before switching the canary on, rather than setting them once the numbers are in. Commercial tools take the same approach: LaunchDarkly's guarded rollouts for AI Configs let you attach guardrail metrics and automatically roll back to the previous variation.

## Step 5: How long should each step run?

Why start at 5% rather than 1%? The tianpan.co piece argues that for LLM features, a 1% canary may never reach the tail of the input distribution, the rare cases where the model performs worse.

The author also notes that an LLM canary may need to run for hours or even days to collect enough labelled or implicitly labelled data.

Go back to the hypothetical scenario at the start. Suppose the client's system handles 2,000 conversations a day, and refund questions make up 2% of traffic. At 5%, the canary receives 100 conversations a day, of which only about 2 concern refunds.

With `MIN_REFUND = 20`, the 5% step needs around 10 days before `decide()` can return "promote". At 1%, the figure is 0.4 refund questions a day, which is close to seeing nothing.

So the duration of each step should be set by the number of samples you need, not by hours. If 10 days is too long, that is a conversation to have with the client at the outset.

| Step | Who gets the canary | Condition to move on (suggested) |
|---|---|---|
| Internal | The client's QA team | No serious errors found on manual review |
| 5% | A small share of real users | Enough samples for rare question types, no guardrail breached |
| 20% | Broader, with clearer comparison data | Quality metrics no worse than control |
| 50% | A/B test, two groups of roughly equal size | Client signs off on the results |
| 100% | All users | Keep the flag for a while longer so you can still roll back |

Flagsmith names only the 5%, 50% and 100% steps. The 20% step in the table is an added buffer that gives you clearer comparison data before running the A/B test at 50%.

## Step 6: Rehearse the kill switch

Flagsmith writes that as soon as they spot a problem, they can turn the flag off instantly. With the code above, setting `FLAG["enabled"] = False` sends every user back to control. Try this once in staging with the client's engineers, and record how many seconds it takes for the change to take effect.

## Three common mistakes

The first is changing prompt and model under the same flag, so when results turn bad you do not know which to revert. The second is logging only the canary group, which leaves no control data when it is time to compare.

The third is jumping from 5% to 50% after a single afternoon because the dashboard looks fine, while the rare questions have not appeared even once. The "wait" branch in `decide()` exists precisely to prevent this.

## How this skill shows up on the job

A request like "can you tweak the prompt so the bot is more polite" is exactly the kind of change this process is built for. Before taking the work on, have answers ready to two questions: if it goes wrong, how do you turn it off, and how long does that take?

When reading job descriptions, look for phrases such as progressive delivery, feature flags, A/B testing or LLM evaluation. On your CV, do not just write "used LaunchDarkly". Be specific: you separated prompt_version into the logs, set guardrail thresholds and minimum sample sizes, took a change to 100% through four steps, and had a kill switch that took effect within a stated number of seconds.

Next time a client asks you to swap the model, reply with a staged rollout schedule, not just a note saying it is done.

**Try this week:**

- Run bucket() on 10,000 fake user_ids to check that the canary share is close to the figure set in the flag and that each user always lands in the same variation
- Add variation and prompt_version fields to the logs of an LLM feature you are working on, then write a query comparing latency between the two groups
- Pick at most 5 guardrail metrics for that feature, and write down the rollback thresholds and minimum sample size for each rare question type before switching the canary on

## Sources

- [Chapter 16 - Canarying Releases (Google SRE Workbook)](https://sre.google/workbook/canarying-releases/)

- [Progressive Delivery for Building LLM-Powered Features](https://flagsmith.com/blog/progressive-delivery-llm-powered-features)

- [Guarded Rollouts for AI Configs](https://launchdarkly.com/changelog/guarded-rollouts-for-ai-configs/)

- [Why Gradual Rollouts Don't Work for AI Features (And What to Do Instead)](https://tianpan.co/blog/2026-04-15-gradual-rollouts-fail-ai-features)

- [LLM Guardrails With Feature Flags: Route, Compare, and Roll Back Safely](https://www.featbit.co/blogs/llm-guardrails-feature-flags)
