Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

Serverless does not mean your application has no state. It means each function execution should be replaceable: a new invocation must be able to work from its event and durable dependencies, without assuming that an earlier invocation left data in the same process. Store business facts in durable data services, and use a durable workflow when the process itself must remember progress across waits, retries, and failures.

What “stateless” means in serverless

A standard function may run in a newly created or previously used environment. Warm reuse can avoid repeated initialization, but it is an optimization rather than a persistence contract. Amazon Web Services states: “For standard Lambda functions, you should assume that the environment exists only for a single invocation.” See AWS’s Designing Lambda applications guidance.

Global variables, temporary files, and in-memory caches can therefore improve one invocation, but they cannot be the authoritative copy of customer, payment, inventory, or job data. A later invocation may run elsewhere, run concurrently, or arrive after the original environment has disappeared.

Separate the two kinds of state

Business state

Business state is the durable record of what your system knows: an order, account balance, uploaded object, document status, or an idempotency key. Write it to a service designed to retain and retrieve data. AWS lists services such as Amazon S3, DynamoDB, and SQS as examples used with Lambda applications. Pass a record identifier or event context to later steps rather than relying on process memory.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

Workflow state

Workflow state is the progress of a multi-step process: which step completed, what is waiting, when to retry, and how to resume after failure. A workflow engine checkpoints that progress and coordinates waits and recovery. It complements a domain database; its execution history is not automatically a substitute for your application’s system of record.

A reliable serverless state pattern

  1. Accept an event and validate it. Extract a stable business identifier and reject malformed input before starting side effects.
  2. Persist the business record. Create or update the order, job, or document in durable storage, including a status that represents the next required action.
  3. Perform one bounded step. A function should read the needed record, do its work, and write the result or a new status.
  4. Record progress outside the runtime. Store checkpoints in the workflow service, or in durable records when the process is simple enough not to need a workflow engine.
  5. Make side effects idempotent. Retries must not charge a card twice, create duplicate shipments, or apply the same update repeatedly. Use an idempotency key, conditional write, or deduplication record appropriate to the operation.
  6. Design for replay and recovery. Assume a step can time out after completing its external action but before reporting success. A retry should safely determine whether the action already happened.

When a workflow service is warranted

Chaining functions manually through ad hoc callbacks, queues, and flags can tightly couple components and leave recovery logic scattered through application code. AWS recommends purpose-built orchestration for complex workflows and distinguishes code-first durable functions from the separately represented, cross-service role of Step Functions. The relevant AWS comparison is documented in Durable functions or Step Functions.

Use a workflow boundary when a process has several dependent steps, long waits, human or external callbacks, retries with different policies, compensation paths, or a need for operators to see exactly where an execution stopped. Keep the domain data in its own durable model even when the workflow engine records execution progress.

Provider examples and the questions they answer

Option Documented capability Decision question
AWS Lambda plus durable services Stateless function design with durable writes to services including S3, DynamoDB, and SQS. Is this state business data that belongs in a database, object store, or queue?
AWS Lambda durable functions Code-first orchestration for Lambda-centric workflows, with checkpointing and recovery. Should workflow logic stay alongside application code?
AWS Step Functions Visually modeled orchestration and coordination across AWS services. Would a separately represented, cross-service workflow improve review and operations?
Azure Durable Functions An Azure Functions extension with orchestrator, activity, and entity functions; the runtime manages state, checkpoints, retries, and recovery. See Microsoft’s overview. Does the application already use Azure Functions, and does this programming model fit the process?
Google Cloud Workflows Managed workflows that can hold state, retry, poll, and wait. Google’s overview documents workflows doing so for up to one year; that is a service limit, not an industry statistic. See Google Cloud Workflows overview. Is a managed sequence of service operations the right boundary?

These are capability sketches, not a ranking of cost, latency, throughput, portability, or security. Current regional limits and service behavior should be checked in the provider documentation for your deployment.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

How to choose where state lives

Put it in a data store when…

  • The value represents a business fact that must outlive one execution.
  • Many requests, services, or reports need to query it.
  • You need domain-specific constraints, indexes, retention, or audit history.

Put progress in a workflow when…

  • The process has ordered steps, waits, timers, retries, or branches.
  • An operator must inspect, resume, or cancel an execution.
  • Recovery logic would otherwise be duplicated across several functions.

Use both when…

  • A durable record is the source of truth and the workflow coordinates how it changes.
  • The workflow may be replayed, while business updates must remain idempotent and auditable.

Failure modes to test before production

  • Duplicate delivery: send the same event twice and verify one business outcome.
  • Timeout after success: make an external call succeed, then force the function to fail before acknowledgment.
  • Cold start or replacement: run each invocation as if no global memory or local file existed.
  • Concurrent updates: invoke two workers for one record and verify conditional writes or version checks prevent corruption.
  • Paused execution: resume a workflow after a long wait and confirm credentials, schemas, and referenced records are still valid.
  • Partial workflow failure: inspect whether the failed step, retry policy, and operator action are visible without reconstructing events from scattered logs.
Independent reader supportYour contribution helps us test, update, and keep practical guides available for everyone.Support on Ko-Fi

The practical rule

Treat every function environment as disposable. Treat durable business data as the source of truth. Treat workflow progress as a separate concern that deserves a managed orchestration mechanism when the process can pause, retry, branch, or recover. This explicit split preserves serverless replaceability without pretending that real applications can avoid state.

Product prices and availability are accurate as of the date/time indicated and are subject to change. Any price and availability information displayed on Amazon at the time of purchase will apply.