The saga pattern in practice: when an agent fails halfway across several systems
Your agent has reserved stock in the warehouse and raised a document in the ERP when a third system refuses the request. A rollback command will not help you at that point. A log that records how to undo every completed step will.
In brief
- A compensating transaction is a new transaction that runs after the original step has committed, and its logic depends on the business process.
- Log each step together with how to undo it, attach an idempotency key to every call, retry first, and compensate only when the workflow cannot go forward.
- Irreversible steps such as sending an email must come after every check. If compensation itself fails, stop and call a person.
- 1reserve_stock: doneStock is held; the log also records the undo command release_stock
- 2create_credit_note: failedRetries exhausted, orchestrator switches to compensation; the email is never sent
- 3reserve_stock: needs_humanrelease_stock also fails; the saga halts and run_saga returns needs_human
- 4reserve_stock: undoneAfter the fix, compensate reruns from the same log; only the unfinished step is undone
Reading the log alone shows which steps ran, which were compensated and which await a human.
Graphic: FDE Times
Picture an agent handling a customer refund request. It reserves the goods in the inventory system, creates a credit note in the ERP, then sends a confirmation email. Step two returns an error, but step one has already committed. The inventory system is now holding a batch of stock that nobody needs.
No transaction spans all three systems, so there is nothing to roll back. Sagas exist for exactly this situation: according to the Azure Architecture Center, a large operation is split into a sequence of local transactions, and when one step fails the saga runs compensating transactions to reverse the steps that came before it.
For a forward deployed engineer this is daily work, not distributed-systems theory. The more write access an agent has to a customer’s systems, the earlier someone in the review will ask what happens if it fails halfway. This guide builds a small Python orchestrator so you have a concrete answer.
What will you build, and what do you need?
The result is an orchestrator that runs three steps in sequence. It logs each step, retries on transient errors, compensates from the log when it cannot go forward, and flags the saga for human review when compensation itself fails. All you need is Python 3 and an editor.
Every external system here is a simulated function. The code has been simplified to teach the principle and is not ready for production.
Step 1: Which steps cannot be reversed?
Before writing any code, classify each step. Microsoft’s saga documentation calls a step that cannot be reversed a pivot transaction: once the pivot succeeds, compensation no longer makes sense, and the steps after it must be retryable until they finish.
An email to the customer is the textbook case, because nobody can recall it. The design rule that follows is clear: an irreversible step runs only once every important check has passed.
STEPS = [
{"name": "reserve_stock", "kind": "compensable", "undo": "release_stock"},
{"name": "create_credit_note", "kind": "compensable", "undo": "void_credit_note"},
{"name": "send_customer_email", "kind": "pivot", "undo": None},
]
Check: every step whose undo is None must sit at the end of the list. If the agent has to send an email midway, that is a sign the process should be split into two workflows, each ending with its own irreversible step.
Step 2: Record how to undo each step as you go
The core mechanism is to record information about each step along with how to undo it. When the operation fails, the workflow walks back through the completed steps. The log must live outside the process’s memory and be reread on startup; otherwise a single crash wipes out all progress.
import json, os
class SagaLog:
def __init__(self, path):
self.path = path
self.entries = []
if os.path.exists(path): # rerun after a crash: load earlier progress
with open(path) as f:
self.entries = json.load(f)
def record(self, step, status, undo=None):
self.entries.append({"step": step, "status": status, "undo": undo})
with open(self.path, "w") as f:
json.dump(self.entries, f, indent=2)
def is_undone(self, step):
return any(e["step"] == step and e["status"] == "undone" for e in self.entries)
Check: call record, kill the process, then create a new SagaLog with the same path. entries must contain exactly the lines you wrote, each showing which step finished and which command will reverse it.
Step 3: Does a duplicate call corrupt data?
Retrying means a step may run more than once, so every step must be an idempotent command. Stripe is a real-world example: for each idempotency key, Stripe stores the status code and body of the first request, whether it succeeded or failed, and returns that same result for every retry that uses the same key.
def idem_key(saga_id, action):
return f"{saga_id}:{action}"
_seen = {} # simulates the customer's system storing results by key
def fake_reserve_stock(key):
if key not in _seen:
_seen[key] = {"reserved": 5}
return _seen[key]
Check: call fake_reserve_stock twice with the same key. The stock must be reserved only once.
Step 4: Retry first, compensate later
A passing network glitch is no reason to cancel the whole transaction. Compensate only when the workflow cannot go forward: when retries are exhausted or the error is known not to be transient.
class Transient(Exception): pass
class Permanent(Exception): pass
def run_with_retry(fn, key, attempts=3):
for _ in range(attempts):
try:
return fn(key)
except Transient:
continue
raise Permanent(f"retries exhausted: {key}")
def run_saga(saga_id, steps, systems, log, priority=()):
done = []
for s in steps:
try:
run_with_retry(systems[s["name"]], idem_key(saga_id, s["name"]))
log.record(s["name"], "done", s["undo"])
done.append(s)
except Permanent:
log.record(s["name"], "failed")
# returns "compensated" or "needs_human" depending on the compensation outcome
return compensate(saga_id, done, systems, log, priority)
return "completed"
This is the orchestration style described in Microsoft’s saga documentation: the orchestrator sends requests, stores and interprets the state of each task, and handles recovery through compensating transactions. The cost is that the orchestrator becomes a potential point of failure, which is one more reason the log must live on disk.
Step 5: In what order do you compensate, and what if compensation fails?
The default is to reverse the order in which steps ran, but that is not mandatory. The data store most sensitive to inconsistency should be undone first. In this example, the accounting ERP probably matters more than the warehouse.
Compensation can fail too. The system therefore needs to record progress so it can resume from exactly the point of failure. Sometimes it has to stop, hand over to a person and send an alert that states the reason.
def compensate(saga_id, done, systems, log, priority=()):
ordered = sorted(reversed(done), key=lambda s: s["name"] not in priority)
for s in ordered:
if log.is_undone(s["name"]):
continue # already compensated in an earlier run, skip
try:
run_with_retry(systems[s["undo"]], idem_key(saga_id, s["undo"]))
log.record(s["name"], "undone")
except Permanent as e:
log.record(s["name"], "needs_human", str(e)) # reason stored in the undo field for brevity
return "needs_human" # stop and wait for human review
return "compensated"
The return status is easy to overlook. If run_saga still reports “compensated” when compensation stopped halfway, the caller will assume everything is clean while the warehouse is still holding stock. That is why needs_human must be a separate status that the caller can check and use to fire an alert.
Check: make create_credit_note raise Permanent. The log must show reserve_stock done, create_credit_note failed, then reserve_stock undone; run_saga must return "compensated", and the email must never be sent. Then make release_stock fail as well: the log must stop at needs_human and run_saga must return "needs_human". Fix release_stock and call compensate again with a fresh SagaLog opened from the same file: any step already marked undone is skipped, and only the unfinished step is compensated.
Step 6: Two sagas touching the same customer
Suppose a customer sends two requests in quick succession. If two sagas run in parallel on the same order, a compensating step from one can slip in between the forward steps of the other. The Sequential Convoy pattern handles this by grouping messages by a key such as the order ID. Each group is processed in sequence, while different groups still run in parallel.
The trap in this approach is the poison message. A message that fails repeatedly blocks every message behind it in the same session. You need to count delivery attempts and move the message to a dead-letter queue once it exceeds the retry threshold. This is not in the sample code, but do not skip it in production.
Common mistakes
The most common mistake is treating compensation as an Undo button. Wikipedia defines a compensating transaction as a new transaction that reverses the effects of an already committed transaction, which makes it different from a rollback.
Compensation also does not return the system to its original state. It has to account for work running concurrently, and that logic depends on the application.
The second mistake is forgetting that sagas lack isolation. The result of the original step is visible to other systems before it is compensated, so lost updates or dirty reads can occur. There are a few familiar countermeasures: a semantic lock (an application-level flag indicating that a record is being updated), commutative updates, and rereading a value before writing it.
The third mistake comes from the agent. For ambiguous or high-impact cases, the workflow should pause for human review. Do not let the LLM decide whether to compensate. Let the orchestrator decide by rule, and let the agent only make suggestions.
What does this skill look like on a customer site?
In the first week, do not write the orchestrator straight away. Sit down with the customer’s operations team and ask about every API the agent will write to: does it accept an idempotency key, what is the undo command, and who is authorised to step in when the undo fails? That table of answers is your STEPS.
When reading FDE job descriptions, look for phrases such as “multi-system workflows”, “ERP/CRM integration” or “agents with write access”. On a CV, a line like “designed a saga for a refund agent: isolated the pivot step, idempotency keys on every write, an escalation process when compensation fails” says more than any claim of “experience with distributed systems”.
An agent will fail halfway at some point. What the customer remembers after the incident is whether you could immediately open a log file showing which steps ran, which were compensated and which are waiting for human review.
Was this article useful?
Thanks for the feedback!
5 sources
- Compensating Transaction Pattern - Azure Architecture Center | Microsoft Learn · 2026-04-16
- Saga Design Pattern - Azure Architecture Center | Microsoft Learn · 2025-02-25
- Sequential Convoy Pattern - Azure Architecture Center | Microsoft Learn · 2026-06-24
- Idempotent requests | Stripe API Reference
- Compensating transaction - Wikipedia