Three replicas, one cron job: how to stop customers getting three emails
Scale a service from one instance to three and its cron job runs three times. Scheduler-Agent-Supervisor and leader election help keep this under control, but the most durable fix isn't a lock.
In brief
- Scaling out adds schedulers, and a run that outlasts its interval also causes overlap, so every scheduled job must be idempotent.
- The Scheduler records each step's state with a complete-by deadline. The Supervisor looks for steps still marked 'running' past that deadline, but only requests recovery and never performs it.
- Leader election is just one option. Conditional claims, distributed locks or CronJob Forbid are often enough, but none of them guarantees exactly-once.
- 1Scheduler claims and logs stepConditional claim via LockedBy; writes status and complete-by deadline to state store
- 2Agent executes the stepDoes the step's work; logic must be idempotent since the step may rerun
- 3Supervisor scans periodicallyFinds steps still 'running' but past their complete-by deadline
- 4Recovery is requestedSupervisor only requests; Scheduler and Agent rerun the step
- ↻ Repeat from step 1
The Supervisor only detects failures and requests recovery, so every step that gets rerun must be idempotent.
Graphic: FDE Times
Suppose you deploy a service that syncs invoices for a customer. The code includes a cron job that runs at 2am. In the first week there is a single replica and everything is fine. In the second week the customer’s infrastructure team scales up to three replicas to handle more load, and the next morning the accounts department calls because every customer has received three payment reminders.
Nobody wrote careless code. Microsoft’s best-practices guidance on background jobs is explicit: scaling out the host that runs the scheduler produces more schedulers, and those schedulers can launch multiple instances of the same task.
There is a second route to duplicate runs, and fewer people notice it. If a run takes longer than the interval between two triggers, the scheduler starts a new instance while the old one is still going.
If the solution you deploy has scheduled jobs, such as data syncs or report delivery, and can run as more than one instance, plan for this from the start. This article shows how to design jobs so the results stay correct however many copies run in parallel.
A lock is the second layer; idempotency is the first
Many engineers’ first instinct is to add a lock so that only one instance runs. Locks are necessary, but the priority runs the other way. Microsoft’s background-job guidance requires scheduled tasks to be designed as idempotent, so that running the same task several times does not produce duplicate results.
The Scheduler-Agent-Supervisor pattern documentation likewise requires each step’s logic to be idempotent, because a step can run more than once when it is retried.
The reason is practical. Kubernetes admits in its CronJob documentation that in some situations a single CronJob can still create several Jobs running at once. When the platform itself concedes that the lock sometimes slips, your code has to survive the slip.
Three roles, one state table
When a job has several steps, such as pulling data from the ERP, calculating outstanding balances and then sending emails, idempotency alone is not enough. You also need to know which step has stalled and who will clean it up. The Scheduler-Agent-Supervisor pattern splits the work across three logical roles so that the whole task succeeds or fails as a single unit.
The Scheduler writes the state of each step to a durable state store, including a complete-by deadline that limits how long the step may run. The Agent executes the step. The Supervisor runs periodically, finds steps that are overdue or have failed, and requests recovery. The Supervisor does not carry out the recovery itself: it only raises the request, and the Scheduler and Agent do the work.
The way the pattern detects failure is also simple. A step that timed out and a step that crashed look identical in the state store: the record still says running, but complete-by has passed. The Supervisor only needs to scan for that condition, without guessing why the Agent died.
A worked example: a billing job on three replicas
Back to the invoicing service. Instead of having cron call the email function directly, you create a job_steps table with the columns step_id, status, locked_by and complete_by. At 2am the Scheduler on all three replicas wakes up, and all three run the following claim statement:
UPDATE job_steps
SET status = 'running',
locked_by = :instance_id,
complete_by = now() + interval '15 minutes'
WHERE step_id = :step_id
AND status = 'pending';
-- rows_affected = 1: this instance claimed the step
-- rows_affected = 0: another instance already took it, exit
(The comments read: one affected row means this instance won the step; zero means another instance already took it, so exit.)
This is a conditional state transition, like Microsoft’s example using a LockedBy field. Several Scheduler orchestrations can run at once, but only one attempt wins the order. The database handles the contention, so there is no need to build a separate election system.
The Supervisor is a small job that runs every few minutes:
SELECT step_id, locked_by FROM job_steps
WHERE status = 'running' AND complete_by < now();
For each row returned, the Supervisor records a recovery request, for example by resetting the step to pending or by publishing a message. The Scheduler picks up the request and hands the step back to an Agent. Because the step may run again, the email send must be idempotent.
The safe approach is to claim the right to send before sending. Create a sent_reminders table with a unique constraint on (customer_id, billing_period) (the second column being the billing period) and insert a row before calling the email service. Whichever instance inserts successfully sends the email; any instance that hits a unique-constraint violation skips it, so two parallel instances cannot both send.
Do not do it the other way round, as “check, send, then record”: two instances can both pass the check before either has written anything, and the customer still gets two emails.
The cost of inserting first is that if the Agent dies just after the insert, the email may never go out. Make this trade-off explicit to the customer, or add a sending state so the Supervisor can find abandoned sends.
When do you actually need leader election?
The Supervisor can also run as several instances. Microsoft’s documentation notes that when multiple Supervisors are active, they must coordinate so they don’t all rush to recover the same failed step. Leader election is one way to do this: elect one instance as leader and let it coordinate the others.
Electing a leader is harder than it looks. The election process must prevent two instances from becoming leader at the same time, and the system must detect when the leader dies, through heartbeats or polling. With a blob lease approach, the lease must have an expiry so a failed leader cannot hold it forever, and the leader must keep renewing it.
That is where the classic trap lies. If the leader’s task hangs while its lease-renewal thread keeps running, the leader renews indefinitely and no other instance can take the lease. So check the health of the work the leader is actually doing, not just whether the lease is still alive.
Microsoft also says plainly that leader election is often unnecessary. A singleton process that is shut down and restarted when it fails, or a simple locking mechanism, is enough. But a shared mutex service also becomes a single point of failure. The table below summarises the common options:
| Mechanism | What it guarantees | Weakness to remember |
|---|---|---|
| Azure Functions timer trigger | Distributed lock so only one instance runs | The job still needs to be idempotent |
CronJob with concurrencyPolicy: Forbid |
Skips a new run if the previous one hasn’t finished | Can still occasionally create several Jobs at once |
Conditional claim (LockedBy) |
Only one attempt wins each step | Depends on the state store |
| Lease-based leader election | One instance coordinates the others | A hung leader can still renew the lease |
Steps to take on a customer project
Start with an inventory. For each scheduled job, ask the customer: what happens if it runs twice, how long does the longest run take relative to the interval, and do the replicas autoscale? The answer on run time gives you the figure for complete-by, and tells you whether the risk of overlapping runs is real.
Then work in order: make each step idempotent, add a state table with complete-by, use conditional claims, and write a Supervisor that scans for overdue steps. Only at that point consider leader election, and only if there is a genuine coordinating role that cannot be split up per record.
Common mistakes
The most common mistake is trusting a single setting, such as setting Forbid and considering the job done, when Kubernetes itself says it doesn’t block every case.
The second is letting the Supervisor rerun steps itself, which blurs the boundary between roles and creates yet another source of duplicate runs. The third is setting complete-by shorter than the real run time, so the Supervisor triggers recovery for steps that are still running normally.
How to talk about this skill when job hunting
On your CV, rather than writing “used cron”, describe how you made jobs idempotent and resilient to multiple replicas, with the number of duplicate runs before and after the fix.
The best test is to cause the failure yourself. Start two instances of the billing job at once, kill one midway, and check whether the system recovers on its own while each customer still receives exactly one email.
Was this article useful?
Thanks for the feedback!