Recommended Free Tools
iTechGuides is reader-supported. When you buy through links on our site, we may earn an affiliate commission. As an Amazon Associate I earn from qualifying purchases. Learn more
A saga does not automatically roll back a business process across services. Each service commits its own local transaction; if a later step fails, the workflow must retry, take a valid alternate path, or run separately designed compensating transactions for completed work. Those actions can move the process toward a valid business state, but they do not guarantee an exact restoration of the state that existed before the saga began. Microsoft Azure Architecture Center and AWS Prescriptive Guidance describe these mechanics as application-level recovery, not cross-service ACID rollback.
Saga rollback mechanics and failure atomicity
A saga coordinates a sequence of local transactions, each committed by a participant service. Coordination can happen through events or through an orchestrator that directs the participants. Because the services do not share one atomic transaction, a failure in a later step does not undo earlier commits. Recovery is another workflow, implemented with application-specific actions. Microsoft’s saga guidance describes local transactions and coordination; Microservices.io’s saga reference also emphasizes that participants update their own data.
Failure atomicity therefore has a narrower meaning in a saga than in a single database transaction. A participant’s local transaction may be atomic within its own service, but the business process as a whole can be partially complete while recovery is pending. Compensation is not a database inverse or a guarantee of restoring a global snapshot. It is a domain operation intended to counteract an earlier effect while respecting current business rules and state.
As the Microsoft Azure Architecture Center puts it: “A compensating transaction doesn’t necessarily return the system data to its state at the start of the original operation.” (Compensating Transaction pattern.)
#1 Best Overall
Forward steps, pivots, and recovery
Plan the forward workflow by distinguishing work that can be compensated from the step that commits the business process to its intended direction and any later retryable work. Azure’s saga guidance uses the terms compensable, pivot, and retryable steps for this distinction. The exact classification depends on the domain: a payment authorization, shipment dispatch, or reservation may have different reversibility and business consequences in different systems. Microsoft Azure Architecture Center
The partial execution trap
The trap is to treat a later failure as though it erased earlier commits. Suppose order creation succeeds and inventory is reserved, but payment authorization fails. The order and reservation already exist in their respective services. Until the workflow resolves the failure, the system is in an intermediate state: some participants reflect the attempted purchase and others do not. This is expected in an eventually consistent design; it becomes a correctness incident if the workflow loses the information required to recover, repeats an unsafe action, ignores concurrent changes, or records compensation as complete when it failed. AWS uses order, inventory, and payment in its saga examples. AWS Prescriptive Guidance
Rank #2
If the payment problem is transient, retrying authorization may allow the workflow to continue. If payment is invalid or forward progress cannot be restored, the domain may call for releasing inventory and canceling or amending the order. A fallback payment method, product substitution, customer choice, or human review may be more appropriate than immediate unwind. These are business decisions, not universal saga rules. Microsoft’s compensating-transaction guidance discusses alternate actions and human intervention; AWS distinguishes continuation from compensation.
What to persist for recovery
Recovery must survive process restarts and partial failures. Persist the saga’s identity, each participant step’s outcome, the context needed to compensate it, and separate progress for compensation attempts. Correlate the forward and recovery activity so operators can tell which business process is stuck and which action remains. Make participant operations safe to repeat where possible, and define alerting and manual intervention for cases that do not converge automatically. Microsoft’s compensating-transaction guidance specifically covers execution state, compensation metadata, retries, monitoring, and escalation. Microsoft Azure Architecture Center
Rank #3
How to choose compensating transaction ordering
Start from the dependency graph and business invariants, not the slogan “undo in reverse.” Record which forward effects depend on others, which are externally visible, which can be repeated, and which are difficult or impossible to reverse. Then choose the order that reduces the risk of leaving participants in an invalid or harmful state. Reverse forward order is a useful starting point for dependent steps, but exact reverse order is not mandatory: some compensations can run in parallel, and a particularly sensitive data store may need to be corrected first. Microsoft Azure Architecture Center
Use domain corrections, not stale snapshots
A compensation should use retained context about the original operation and apply a valid corrective action. Restoring an old snapshot can overwrite legitimate updates made after the saga began. For example, releasing a reservation should target the reservation created by this saga rather than reset an inventory quantity to a previously observed value. The precise action must account for intervening work and the service’s business rules. Microsoft Azure Architecture Center
Rank #4
Make points of no return explicit
Identify irreversible, externally visible, or legally binding steps before implementation. Where possible, put critical validations before such steps, and define what happens if a later action fails after one has occurred. The outcome might be a corrective business action, a hold for review, or a workflow that continues by another route; it may not be possible to erase the effect. Microsoft Azure Architecture Center
The Tool Desk
Outbyte Driver Updater FREEFix the driver behind crashes, sound loss and screen glitchesFind Drivers →Outbyte PC Repair FREEClear out junk files and repair common Windows errorsFree Scan →Choose retry, compensation, an alternate path, or review
| Condition | Recovery direction | Design consideration |
|---|---|---|
| Temporary infrastructure or network failure | Retry the local transaction and continue forward if progress remains possible. | Participants must tolerate repeated execution; otherwise a retry can duplicate an effect. AWS Prescriptive Guidance; Microsoft Azure Architecture Center |
| Nontransient business failure, such as invalid payment | Compensate completed steps if the process cannot proceed. | Define compensation as a domain operation; it need not be an exact inverse. AWS Prescriptive Guidance; Microsoft Azure Architecture Center |
| A valid replacement service or route exists | Continue through the alternate path when the business outcome permits. | Do not automatically unwind if a domain rule or customer choice should determine the route. Microsoft Azure Architecture Center |
| High-impact or ambiguous outcome | Pause for human review where appropriate. | Preserve state and define how an operator can resume or compensate the workflow. Microsoft Azure Architecture Center |
| A compensating action fails | Track its status, retry safely, alert, and provide a manual intervention path. | The business process can remain inconsistent until recovery succeeds; do not mark the compensation complete prematurely. Microsoft Azure Architecture Center; Microsoft Azure Architecture Center |
Choreography or orchestration?
Both approaches coordinate local transactions; neither creates cross-service ACID isolation. Choose based on workflow complexity, how clearly operators and developers need to see state, and how coordination failure should be handled.
| Approach | How it coordinates | Trade-offs |
|---|---|---|
| Choreography | Participants react to events and publish further events. | Avoids a central coordinator and can suit a small participant set, but the event dependency graph can become difficult to follow as services are added. AWS Prescriptive Guidance; Microsoft Azure Architecture Center |
| Orchestration | A coordinator records or interprets workflow state and directs participants. | Can make a complex flow easier to follow and reduce participant-to-participant dependencies, but adds coordination logic and a potential central failure point. AWS documents Step Functions as one implementation example, not a requirement for sagas. AWS Prescriptive Guidance; Microsoft Azure Architecture Center |
Whichever style you choose, make local state changes and message publication reliable. Microservices.io identifies the transactional outbox and event sourcing among related approaches for this problem. Microservices.io
Concurrency and isolation: what a saga does not protect
A saga coordinates workflow steps but does not provide transaction isolation across participant databases. Concurrent workflows may read stale values or overwrite one another; Azure lists anomalies including lost updates, dirty reads, and fuzzy or nonrepeatable reads. AWS also identifies lack of isolation as an orchestration concern. Microsoft Azure Architecture Center; AWS Prescriptive Guidance
Choose controls that match the invariant at risk. Options described in the guidance include semantic locks, commutative updates, rereading values before updating, and recording or versioning operation order. These controls reduce specific concurrency risks; they do not turn the saga into a distributed ACID transaction. Microsoft Azure Architecture Center; AWS Prescriptive Guidance
Quick Recap
Design review checklist
- For every forward step, document its committed effect, whether it can be repeated, whether it can be compensated, and what context recovery needs.
- Classify transient failures separately from business-rule failures; state when the workflow retries, changes route, compensates, or pauses.
- Specify compensation dependencies and justify any order that differs from reverse forward order; identify safe parallel compensations.
- Define the behavior after an irreversible or externally visible step, including cases where exact restoration is impossible.
- Persist forward and compensation status, correlate activity across participants, and alert on stuck or repeatedly failing recovery.
- Test duplicate delivery, participant timeouts, process restarts, concurrent updates, and compensation failure—not only the successful path.
- Choose concurrency controls around explicit business invariants, and ensure local state changes and message publication cannot silently diverge.
Product prices and availability are accurate as of the date/time indicated and are subject to change. Any price and availability information displayed on Amazon at the time of purchase will apply.

