# Three replicas, one cron job: how to stop customers getting three emails

> Scale a service from one instance to three and its cron job runs three times. Scheduler-Agent-Supervisor and leader election help keep this under control, but the most durable fix isn't a lock.

Original: https://fdetimes.net/en/guides/cron-jobs-multiple-replicas-idempotency/

Suppose you deploy a service that syncs invoices for a customer. The code includes a cron job that runs at 2am. In the first week there is a single replica and everything is fine. In the second week the customer's infrastructure team scales up to three replicas to handle more load, and the next morning the accounts department calls because every customer has received three payment reminders.

Nobody wrote careless code. Microsoft's best-practices guidance on background jobs is explicit: scaling out the host that runs the scheduler produces more schedulers, and those schedulers can launch multiple instances of the same task.

There is a second route to duplicate runs, and fewer people notice it. If a run takes longer than the interval between two triggers, the scheduler starts a new instance while the old one is still going.

If the solution you deploy has scheduled jobs, such as data syncs or report delivery, and can run as more than one instance, plan for this from the start. This article shows how to design jobs so the results stay correct however many copies run in parallel.

## A lock is the second layer; idempotency is the first

Many engineers' first instinct is to add a lock so that only one instance runs. Locks are necessary, but the priority runs the other way. Microsoft's background-job guidance requires scheduled tasks to be designed as idempotent, so that running the same task several times does not produce duplicate results.

The Scheduler-Agent-Supervisor pattern documentation likewise requires each step's logic to be idempotent, because a step can run more than once when it is retried.

The reason is practical. Kubernetes admits in its CronJob documentation that in some situations a single CronJob can still create several Jobs running at once. When the platform itself concedes that the lock sometimes slips, your code has to survive the slip.

**Key point:** Locks reduce the number of duplicate runs; only idempotency makes duplicate runs harmless.

## Three roles, one state table

When a job has several steps, such as pulling data from the ERP, calculating outstanding balances and then sending emails, idempotency alone is not enough. You also need to know which step has stalled and who will clean it up. The Scheduler-Agent-Supervisor pattern splits the work across three logical roles so that the whole task succeeds or fails as a single unit.

The Scheduler writes the state of each step to a durable state store, including a *complete-by* deadline that limits how long the step may run. The Agent executes the step. The Supervisor runs periodically, finds steps that are overdue or have failed, and requests recovery. The Supervisor does not carry out the recovery itself: it only raises the request, and the Scheduler and Agent do the work.

The way the pattern detects failure is also simple. A step that timed out and a step that crashed look identical in the state store: the record still says *running*, but complete-by has passed. The Supervisor only needs to scan for that condition, without guessing why the Agent died.

## A worked example: a billing job on three replicas

Back to the invoicing service. Instead of having cron call the email function directly, you create a `job_steps` table with the columns `step_id`, `status`, `locked_by` and `complete_by`. At 2am the Scheduler on all three replicas wakes up, and all three run the following claim statement:

```sql
UPDATE job_steps
SET status = 'running',
    locked_by = :instance_id,
    complete_by = now() + interval '15 minutes'
WHERE step_id = :step_id
  AND status = 'pending';
-- rows_affected = 1: this instance claimed the step
-- rows_affected = 0: another instance already took it, exit
```

(The comments read: one affected row means this instance won the step; zero means another instance already took it, so exit.)

This is a conditional state transition, like Microsoft's example using a `LockedBy` field. Several Scheduler orchestrations can run at once, but only one attempt wins the order. The database handles the contention, so there is no need to build a separate election system.

The Supervisor is a small job that runs every few minutes:

```sql
SELECT step_id, locked_by FROM job_steps
WHERE status = 'running' AND complete_by < now();
```

For each row returned, the Supervisor records a recovery request, for example by resetting the step to `pending` or by publishing a message. The Scheduler picks up the request and hands the step back to an Agent. Because the step may run again, the email send must be idempotent.

The safe approach is to claim the right to send before sending. Create a `sent_reminders` table with a unique constraint on `(customer_id, billing_period)` (the second column being the billing period) and insert a row before calling the email service. Whichever instance inserts successfully sends the email; any instance that hits a unique-constraint violation skips it, so two parallel instances cannot both send.

Do not do it the other way round, as "check, send, then record": two instances can both pass the check before either has written anything, and the customer still gets two emails.

The cost of inserting first is that if the Agent dies just after the insert, the email may never go out. Make this trade-off explicit to the customer, or add a `sending` state so the Supervisor can find abandoned sends.

## When do you actually need leader election?

The Supervisor can also run as several instances. Microsoft's documentation notes that when multiple Supervisors are active, they must coordinate so they don't all rush to recover the same failed step. Leader election is one way to do this: elect one instance as leader and let it coordinate the others.

Electing a leader is harder than it looks. The election process must prevent two instances from becoming leader at the same time, and the system must detect when the leader dies, through heartbeats or polling. With a blob lease approach, the lease must have an expiry so a failed leader cannot hold it forever, and the leader must keep renewing it.

That is where the classic trap lies. If the leader's task hangs while its lease-renewal thread keeps running, the leader renews indefinitely and no other instance can take the lease. So check the health of the work the leader is actually doing, not just whether the lease is still alive.

Microsoft also says plainly that leader election is often unnecessary. A singleton process that is shut down and restarted when it fails, or a simple locking mechanism, is enough. But a shared mutex service also becomes a single point of failure. The table below summarises the common options:

| Mechanism | What it guarantees | Weakness to remember |
|---|---|---|
| Azure Functions timer trigger | Distributed lock so only one instance runs | The job still needs to be idempotent |
| CronJob with `concurrencyPolicy: Forbid` | Skips a new run if the previous one hasn't finished | Can still occasionally create several Jobs at once |
| Conditional claim (`LockedBy`) | Only one attempt wins each step | Depends on the state store |
| Lease-based leader election | One instance coordinates the others | A hung leader can still renew the lease |

## Steps to take on a customer project

Start with an inventory. For each scheduled job, ask the customer: what happens if it runs twice, how long does the longest run take relative to the interval, and do the replicas autoscale? The answer on run time gives you the figure for complete-by, and tells you whether the risk of overlapping runs is real.

Then work in order: make each step idempotent, add a state table with complete-by, use conditional claims, and write a Supervisor that scans for overdue steps. Only at that point consider leader election, and only if there is a genuine coordinating role that cannot be split up per record.

## Common mistakes

The most common mistake is trusting a single setting, such as setting `Forbid` and considering the job done, when Kubernetes itself says it doesn't block every case.

The second is letting the Supervisor rerun steps itself, which blurs the boundary between roles and creates yet another source of duplicate runs. The third is setting complete-by shorter than the real run time, so the Supervisor triggers recovery for steps that are still running normally.

## How to talk about this skill when job hunting

On your CV, rather than writing "used cron", describe how you made jobs idempotent and resilient to multiple replicas, with the number of duplicate runs before and after the fix.

The best test is to cause the failure yourself. Start two instances of the billing job at once, kill one midway, and check whether the system recovers on its own while each customer still receives exactly one email.

**Try this week:**

- List every scheduled job in your current project and note next to each one: what happens if it runs twice at the same time?
- Add status, locked_by and complete_by columns to one job's state table, write the conditional claim UPDATE, then try running it in parallel from two terminals
- Check the Kubernetes CronJobs you use: what is concurrencyPolicy set to, and would the job still be safe if Kubernetes created two Jobs at once anyway?

## Sources

- [Scheduler Agent Supervisor pattern - Azure Architecture Center | Microsoft Learn](https://learn.microsoft.com/en-us/azure/architecture/patterns/scheduler-agent-supervisor)

- [Leader Election pattern - Azure Architecture Center | Microsoft Learn](https://learn.microsoft.com/en-us/azure/architecture/patterns/leader-election)

- [Best Practices for Background Jobs - Azure Architecture Center | Microsoft Learn](https://learn.microsoft.com/en-us/azure/architecture/best-practices/background-jobs)

- [CronJob | Kubernetes](https://kubernetes.io/docs/concepts/workloads/controllers/cron-jobs/)
