Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

To connect Amazon SQS to an AWS Lambda function, create a queue and a separate dead-letter queue (DLQ), attach a redrive policy to the source queue, then create a Lambda event source mapping for that queue. Lambda polls SQS and invokes the function with batches of messages; SQS visibility timeouts and redrive settings govern what happens when processing fails.

The important design choices are how long a message stays invisible during processing, whether a failed batch is retried as a whole or record by record, and how operators investigate and replay quarantined messages. The Terraform example below targets AWS provider 6.19.0 for the queue resources. Pin your provider and check the documentation for that version before applying.

How SQS, Lambda, retries, and a DLQ fit together

SQS is the event source; Lambda does not receive a direct push from the queue. A Lambda event source mapping polls the queue and invokes the function with one or more messages. The mapping is a separate AWS resource from both the queue and the function.

When Lambda processes a message, SQS keeps it hidden from other consumers for the queue’s visibility timeout. If processing succeeds, Lambda deletes the message. If processing fails or the invocation is throttled, messages can become visible again after the timeout and be delivered again. The source queue’s redrive policy moves a repeatedly received message to its configured DLQ after it reaches the receive threshold.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.
  • Event source mapping: controls how Lambda consumes the queue, including batch size and partial-batch response behavior.
  • Visibility timeout: gives an invocation time to process messages before they can be received again.
  • Redrive policy: specifies the DLQ and the maximum receive count before SQS moves a message out of the source queue.

These are separate mechanisms: the mapping handles invocation and retry behavior, while SQS applies the redrive policy. A DLQ contains messages for investigation or controlled recovery; it does not diagnose or repair the cause of failure.

Choose queue type, batch size, and failure behavior

Standard or FIFO queue

Choose a standard queue when strict ordering is not required. AWS Lambda documentation allows a configured SQS event-source batch size of up to 10,000 records for standard queues and up to 10 for FIFO queues. Those are maximum settings, not promises that every invocation will contain that many messages: the synchronous invocation payload quota is 6 MB, including message metadata, so Lambda may invoke with fewer records.

FIFO queues are for workflows that need ordered processing. Consider the effect of moving a failed message to a DLQ before using one: Amazon SQS warns, “Don’t use a dead-letter queue with a FIFO queue if you don’t want to break the exact order of messages or operations.” A quarantined message may no longer be processed in its original position relative to later messages.

Whole-batch retry or partial batch reporting

By default, if the function reports an error while processing a batch, the batch is treated as failed. Messages that the function already processed successfully can be delivered again along with the failed message. That can add unnecessary work and requires handlers to tolerate duplicate processing.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

With the mapping’s ReportBatchItemFailures response type enabled, a handler can report only the failed SQS message IDs. Lambda can then retry those records rather than treating the entire batch as failed. The handler must return the correct identifiers; enabling the mapping option alone does not implement per-record error handling. AWS also notes that with partial batch reporting enabled, Lambda does not scale down message polling when invocations fail, so this choice affects polling behavior as well as retry precision.

Regardless of batch mode, design processing to tolerate duplicate deliveries. The correct idempotency key and persistence strategy depend on the application’s operation; the queue configuration cannot make a non-idempotent handler safe.

Batch size and batching window

A larger batch can reduce invocation overhead, but it increases the amount of work handled together and can make a batch retry more costly if the handler fails. A batch size over 10 for a standard queue requires a batching window of at least one second. The payload limit can also cause actual batches to be smaller than the configured maximum.

Set the visibility timeout and receive threshold

AWS Lambda recommends setting the source queue’s visibility timeout to at least six times the function timeout. When the event source mapping uses a nonzero batching window, include that window in the calculation: use at least six times the function timeout plus the batching window. This is AWS guidance for allowing time to process batches and handle retries, not a substitute for choosing a realistic function timeout.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

For example, a function configured with a 30-second timeout and no batching window calls for a visibility timeout of at least 180 seconds under that recommendation. If the mapping uses a 2-second window, the corresponding minimum is 182 seconds. AWS rejects creation or update of a mapping if the function timeout is longer than the queue’s visibility timeout.

The SQS SetQueueAttributes API documents a visibility-timeout range of 0–43,200 seconds (12 hours) and a default of 30 seconds. Treat the six-times calculation as a recommended minimum; configure the actual value based on the function timeout, batching window, and recovery behavior you expect.

AWS Lambda recommends a source-queue maxReceiveCount of at least 5 as a starting point. More receives give transient failures additional chances to recover; fewer can quarantine a poison message sooner. Set the threshold according to how long a plausible transient incident may last and how quickly operators can respond. The SQS API documents a default receive count of 10 when the redrive attribute is omitted; that API default is not a production recommendation.

Configure queues and the Lambda event source mapping with Terraform

This example creates standard source and dead-letter queues, associates them with an SQS redrive policy, and connects an existing Lambda function to the source queue. The values assume the function timeout is 30 seconds and the mapping uses no batching window. Replace the example names as needed, and make sure the function and queue are in the same Region; they may be in different AWS accounts if cross-account access is configured.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.
terraform {
  required_providers {
    aws = {
      source  = "hashicorp/aws"
      version = "~> 6.19"
    }
  }
}

resource "aws_sqs_queue" "orders_dlq" {
  name = "orders-dlq"
}

resource "aws_sqs_queue" "orders" {
  name                       = "orders"
  visibility_timeout_seconds = 180
}

resource "aws_sqs_queue_redrive_policy" "orders" {
  queue_url = aws_sqs_queue.orders.id

  redrive_policy = jsonencode({
    deadLetterTargetArn = aws_sqs_queue.orders_dlq.arn
    maxReceiveCount     = 5
  })
}

resource "aws_lambda_event_source_mapping" "orders" {
  event_source_arn        = aws_sqs_queue.orders.arn
  function_name           = var.lambda_function_name
  batch_size              = 10
  function_response_types = ["ReportBatchItemFailures"]
}

The queue redrive policy uses the DLQ ARN and the receive threshold; the event source mapping refers to the source queue ARN. HashiCorp’s current AWS provider guidance prefers the dedicated aws_sqs_queue_redrive_policy resource over inline queue redrive-policy attributes for drift detection. The example pins the provider to the 6.19 minor series; provider behavior and arguments can evolve, so confirm them in the documentation matching the version you lock before applying.

The mapping enables partial-batch reporting, but your handler still needs to process records independently and return failed message IDs in the response. If you prefer whole-batch retries, omit function_response_types and handle the resulting duplicate processing appropriately. If you set a nonzero maximum_batching_window_in_seconds, recalculate the visibility timeout using AWS’s formula.

Independent reader supportYour contribution helps us test, update, and keep practical guides available for everyone.Support on Ko-Fi

Validate execution-role permissions and encryption

The Lambda execution role needs permission to consume from the source queue. AWS documents the managed policy AWSLambdaSQSQueueExecutionRole as including the permissions Lambda needs to read SQS messages. For production, validate that the role’s access is appropriately scoped rather than granting broader access than the function needs.

If the queue is encrypted with a customer-managed KMS key, the execution role also needs kms:Decrypt permission on that key. Queue access and key access are separate authorization checks; a mapping can exist while invocations fail because one of them is missing.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

Monitor and operate the DLQ

A DLQ is a quarantine point, not an automatic recovery system. Establish an operational path to notice when messages arrive, inspect representative failures, determine whether the underlying issue is fixed, and choose whether to replay or discard affected messages.

  • Monitor the DLQ for message growth and alert operators when messages appear or accumulate.
  • Inspect message contents and failure context using the application’s logging and tracing practices, while protecting sensitive data.
  • Define who can authorize replay, how replay is performed, and how duplicate side effects are prevented.
  • Track whether a replay succeeds; otherwise a recurring failure may send the same work back into quarantine.

For FIFO workloads, include ordering consequences in that decision: replaying or moving a message can change the sequence in which operations are processed.

Implementation checks before applying

  1. Confirm the source queue and Lambda function are in the same AWS Region.
  2. Check that the execution role can read the source queue and, when applicable, decrypt its messages with the queue’s KMS key.
  3. Set visibility timeout to at least six times the function timeout, plus the batching window when it is nonzero.
  4. Choose a receive threshold that allows plausible transient recovery; AWS Lambda’s documented starting recommendation is at least five receives.
  5. Decide whether to retry whole batches or enable partial batch reporting, and ensure the handler matches that choice.
  6. Plan DLQ alerting, inspection, and controlled replay before relying on the DLQ for production failures.

Product prices and availability are accurate as of the date/time indicated and are subject to change. Any price and availability information displayed on Amazon at the time of purchase will apply.