Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

Build a long-running workflow to expect failures, save progress at durable boundaries, and make every external action safe to repeat. Start by classifying failures and naming workflow steps; then configure bounded retries, rely on the orchestration engine to persist progress, protect side effects with idempotency or compensation, and alert on terminal failures, timeouts, stalled progress, and growing dead-letter queues.

How do I retry a failed workflow step?

Retry only errors that may clear with time or another attempt. A validation error, missing permission, or rejected business rule usually will not improve through repetition, so route it to a terminal failure or compensation path unless you have a specific reason to retry it.

1. Define steps and failure classes

Give each activity or state a clear name. For every step, define what success means and distinguish transient failure, permanent failure, timeout, and cancellation. Assign a stable workflow or execution ID and carry it into logs, error records, and alerts so an operator can find the same run across systems.

2. Set a bounded retry policy per step

Choose which errors match the policy, how many retries are allowed, and how long to wait between attempts. Use capped backoff appropriate to the downstream service; when the engine supports it, jitter can help avoid synchronized retry spikes. Check the selected engine and version for the exact available options and limits rather than assuming every platform supports the same controls.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

For AWS Step Functions, Task, Parallel, and Map states can use ordered Retry and Catch rules. A retrier can specify matched errors with ErrorEquals, an initial delay with IntervalSeconds, a retry limit with MaxAttempts, and exponential growth with BackoffRate. A catcher can route an exhausted or otherwise unmatched failure to an explicit state. See AWS’s Step Functions error-handling guide for configuration and redrive behavior.

After the retry limit is exhausted, make the result visible: transition to a failure-handling path, preserve the failure context, and set the workflow status to something operators can recognize. Do not treat a retried step attempt as proof that the whole workflow has failed—or as proof that it will eventually succeed.

How can a long-running workflow resume after a worker restart?

Use the workflow runtime’s durable progress mechanism rather than keeping essential state only in worker memory. The boundary and recovery model vary by engine: some runtimes checkpoint when orchestration code yields, while others rebuild workflow state by replaying recorded history.

Use durable boundaries in the orchestration runtime

In Azure Durable Task, an orchestrator checkpoints when it yields at an await or yield boundary. Microsoft’s Durable Orchestrations overview describes long-running orchestrations and retry policies for activity and sub-orchestrator calls. Treat calls across these boundaries as the points where you need to reason about persisted progress and recovery.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.
Rank #2
Sale
PowerShell for Sysadmins: Workflow Automation Made Easy
  • Book - powershell for sysadmins: workflow automation made easy
  • Language: english
  • Binding: paperback

Temporal takes a different approach: it persists event history and replays workflow code to reconstruct state after worker loss. Its task documentation explains that replay makes workflows durable and fault-tolerant. Workflow code and its interactions with activities therefore need to follow the runtime’s replay constraints.

Do not confuse a checkpoint with an atomic external action

A persisted workflow position does not make a payment, database write, message send, or other external side effect atomic with the checkpoint. A crash or timeout can leave the outcome uncertain: the external system may have completed the action even though the workflow did not record that completion. A retry or replay can then attempt the action again.

  • Send an idempotency key derived from the workflow and operation to downstream services that support it.
  • Where the downstream service has no idempotency support, keep a durable deduplication record and check it before applying the effect again.
  • If an action cannot safely be repeated or reversed, define a compensating action or route the case for operator review.

These are design safeguards for retry and replay behavior, not guarantees provided by the workflow engines.

Distinguish a failed step from a failed workflow

In Temporal, workflow task failures are automatically retried, while a business-level workflow execution failure needs a configured retry policy to run again. Activity retries and workflow execution retries are separate decisions; choose the scope deliberately and inspect Temporal’s task behavior documentation when configuring them.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

How do I get alerted when a workflow fails or times out?

Alert on conditions that require a person to act, not on every expected wait or retry. A useful notification identifies the workflow, what failed, how long it has been running, and where to inspect its history.

Choose alert conditions that match workflow behavior

  • Terminal failure: the execution has exhausted its retry or recovery path.
  • Timeout: a step or whole execution exceeded its defined limit.
  • Stalled progress: the workflow has not advanced within an age threshold that accounts for normal timers and external waits.
  • Dead-letter queue growth: asynchronous failures are accumulating rather than being safely reprocessed.

For monitor or polling workflows, alert on missing progress relative to the expected polling interval and overall timeout. Avoid treating a deliberate wait as a failure. Azure’s monitor pattern supports waiting between checks, adjusting intervals, and ending when a condition or timeout is reached.

Put actionable context in the notification

Include the workflow or execution ID, failed step or activity, failure class, attempt number, last progress time, and a link or clear method for opening the run history. Emit structured logs with the execution ID and step name so operators can correlate the alert with the workflow’s events. AWS’s Lambda durable functions best practices recommend structured logging, CloudWatch alarms, tracing, terminal-state notifications through EventBridge, and monitoring dead-letter queue depth.

For asynchronous AWS Lambda durable executions, do not assume standard Lambda retry behavior applies: AWS says the service does not automatically retry durable executions on failure. Its guidance recommends preserving terminal failures in a dead-letter queue and monitoring queue depth as well as execution status. EventBridge notifications can cover FAILED, STOPPED, and TIMED_OUT changes.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

Use the diagnostics for your hosting backend

Azure’s diagnostics differ by backend. For Durable Task Scheduler, Microsoft points operators to the scheduler dashboard, Application Insights, Azure portal traces, and Durable Functions Monitor, depending on the diagnostic need. Follow the backend-appropriate guidance in Microsoft’s Durable Functions diagnostics documentation; avoid building operations around internal storage tables that may change.

Independent reader supportYour contribution helps us test, update, and keep practical guides available for everyone.Support on Ko-Fi

How do retry and checkpoint behavior differ by platform?

The terms retry, checkpoint, and alert do not describe identical guarantees across workflow products. Compare the runtime’s retry scope, persisted-progress model, recovery controls, side-effect handling, and operational tooling before choosing how to implement a workflow.

Platform Retry and failure behavior Progress and operational visibility
AWS Step Functions Task, Parallel, and Map states support ordered retry and catch rules with error matching, interval, attempt limit, and backoff settings. Redrive can reset retry counts for rerun states. Use state history and explicit catcher paths for recovery; confirm current service limits and deployment-specific behavior in the AWS guide.
AWS Lambda durable functions Durable executions are not automatically retried on failure; configure explicit retry strategies for transient faults. AWS recommends CloudWatch alarms, structured logs, tracing, EventBridge terminal-state notifications, and DLQ depth monitoring. See Lambda durable functions best practices.
Azure Durable Functions / Durable Task Retry policies are available for activity and sub-orchestrator calls. Orchestrators checkpoint at await or yield boundaries. Diagnostics depend on the scheduler or storage backend; see the orchestration overview and diagnostics guide.
Temporal Workflow task failures retry automatically; workflow execution failures need a configured retry policy. Activity heartbeats can carry payloads across attempts. Workflow state is reconstructed through event-history replay. See Temporal’s task documentation for replay and task behavior.

Azure’s in-process Durable Functions model has a documented support-end date of November 10, 2026; Microsoft recommends migration to the isolated worker model. Check the current Microsoft lifecycle and migration guidance for the version you deploy.

How do I monitor a long-running job until it completes?

Model monitoring as part of the workflow’s lifecycle rather than as an unrelated stream of overlapping polls. A monitor should wait between checks, evaluate a condition, adapt its interval when appropriate, and finish on either success or an explicit timeout. Set age thresholds using the workflow’s normal waiting periods so a healthy execution is not mistaken for a stalled one.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

Before deploying, exercise recovery paths in the target implementation. Confirm the expected behavior for transient failure followed by success, retry exhaustion, timeout, worker restart or replay, duplicate activity attempt, and alert delivery. Also verify that a dead-lettered event can be inspected and safely replayed or reprocessed without repeating an external effect.

Product prices and availability are accurate as of the date/time indicated and are subject to change. Any price and availability information displayed on Amazon at the time of purchase will apply.