Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

iTechGuides is reader-supported. When you buy through links on our site, we may earn an affiliate commission. As an Amazon Associate I earn from qualifying purchases. Learn more

A poison message in Kafka is a record that keeps failing, not a diagnosis. Classify the failure first: retry transient problems for a bounded time; quarantine malformed or persistently rejected records with enough context to investigate; and decide explicitly whether processing should stop or continue. Before you replay a Kafka message, account for duplicate side effects and coordinate the replay with consumer position.

What makes a Kafka message “poison”?

“Poison message” describes a symptom: a record repeatedly fails processing. The cause might be a temporary downstream outage, invalid serialization, or application-level validation. Those failures need different responses. Repeating a retry cannot repair malformed data, while sending a brief infrastructure interruption straight to a dead-letter queue may be premature.

Start by identifying the failure class and the consequence of each possible response:

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.
Failure or handling choice When it fits Trade-off
Bounded retry A transient failure that may clear, such as a temporarily unavailable dependency. Can recover without operator intervention, but a blocking retry can hold up later records in the processing path.
Stop or fail Processing must not advance past a record that has not been handled correctly. Preserves the record in the normal processing path, but can halt progress until the cause is fixed.
Continue and quarantine The application can safely move on while retaining the failed record for investigation and recovery. Processing advances, but the record is not successfully processed; operators need a defined reconciliation or replay path.
Dead-letter queue (DLQ) A record should leave the main processing path after a chosen failure policy, with context retained for later action. Creates a place to inspect and recover failures; it does not diagnose, repair, or automatically replay them.

These are design choices, not interchangeable Kafka guarantees. In particular, asynchronous or retry-topic patterns can change ordering and require their own scheduling policy. Choose whether later records may overtake a failure based on the application’s ordering requirements.

What should a Kafka DLQ record preserve?

A Kafka DLQ is a topic. It is a holding path for failed records, not an automatic repair mechanism. Preserve enough identity and failure information to determine what failed, where it came from, and what happened during delivery. Useful fields include the source topic, partition, offset, consumer group, delivery count, and failure message.

Consider whether to copy the original key, value, and headers. That can make investigation or replay easier, but it can also duplicate sensitive data and increase storage and processing costs. Kafka Connect warns that including message contents in logs can expose sensitive information. KIP-1191 likewise makes copying original content a configurable choice in its share-group proposal, rather than an assumed requirement.

Set access, retention, and monitoring policies for the DLQ as deliberately as for the source topic. Watch for growing backlog and repeated failure patterns, and preserve original identity when designing a recovery path. Without that context, operators may be unable to distinguish a new failure from a duplicate delivery.

Free tools Windows power users keep installed

One-click scans. No signup required.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

How retries and DLQs differ by Kafka component

Kafka Connect, Kafka Streams, and share groups expose different error-handling mechanisms. Do not assume a setting or broker behavior from one applies to another.

Kafka Connect: configure retries and tolerance deliberately

The Apache Kafka 3.5 Connect User Guide documents fail-fast behavior by default. Its default-equivalent error settings include errors.retry.timeout=0, errors.log.enable=false, no configured DLQ topic, and errors.tolerance=none. Consult the guide matching the deployed Connect version before applying settings.

The guide’s example uses a ten-minute retry budget and a maximum delay of thirty seconds. Those are illustrative configuration values, not universal recommendations:

errors.retry.timeout=600000
errors.retry.delay.max.ms=30000
errors.log.enable=true
errors.log.include.messages=false
errors.deadletterqueue.topic.name=connect-errors
errors.tolerance=all

In that example, errors.retry.timeout sets the retry time budget, errors.retry.delay.max.ms caps the delay, the logging options enable error-context logs without message contents, and the DLQ property names the topic. errors.tolerance=all permits the connector to continue while reporting errors. That is a data-handling decision: continuing can keep work moving, but it does not mean the failed record was processed successfully. Define how it will be inspected and reconciled.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

Connect’s exactly-once support depends on the connector implementation and its ability to use framework capabilities. A Connect setting alone does not establish exactly-once effects in an arbitrary external destination.

Kafka Streams: fail, continue, or route through custom handling

The Kafka 4.2 Streams configuration guide describes deserialization exception handlers that return FAIL or CONTINUE. The built-in log-and-fail behavior stops the pipeline on a deserialization failure. Log-and-continue logs the failure and allows later records to be processed, so the errored record leaves normal processing without becoming a successful result.

A custom handler can instead forward a corrupt record to a quarantine topic. If you choose to continue or quarantine, retain the error context and define how the record will be corrected, replayed, or reconciled. Use the documentation for the Kafka Streams version actually deployed; the cited behavior is from the 4.2 guide.

Share groups: verify broker support for the DLQ behavior

KIP-1191, an accepted Apache Kafka proposal last updated July 16, 2026, describes DLQ behavior for share groups. In the proposal, a consumer can reject a record or the delivery-attempt limit can be reached; the record then moves through an archiving state and is written to a configured DLQ topic. Proposed context headers include source topic, partition, offset, group, delivery count, and failure message. Copying the original key, value, and headers is configurable and false by default in the proposal.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

This is proposal-specific behavior, not a reason to assume every deployed broker supports it. The proposal also requires explicit DLQ enablement, restricts permitted topic names with a configurable prefix (default dlq. in the proposal), and does not create DLQ topics automatically by default. Check the broker release and actual deployment configuration before relying on these controls.

KIP-1191 also notes that writing a DLQ record and updating internal share-group state are not atomic. In a rare case, more than one DLQ record can be written for an undeliverable original. Some DLQ write errors are retried; others are logged while the record proceeds to archived state. Make DLQ consumers tolerant of duplicates rather than assuming one archived record always means one DLQ entry.

Independent reader supportYour contribution helps us test, update, and keep practical guides available for everyone.Support on Ko-Fi

Why replaying a Kafka message can repeat side effects

Replaying a record can execute its processing logic again. Kafka’s 4.0 message-delivery design documentation describes a common cause: a consumer processes a record, then crashes before saving its position. A replacement consumer can receive the already-processed record again. This at-least-once delivery behavior can duplicate effects such as charging an account, sending a notification, or writing to a database.

For Kafka-to-Kafka processing, Kafka transactions can atomically commit output records with the input position. That coordinates Kafka input and output; it does not make arbitrary external database or API effects exactly once. For an external destination, use destination cooperation or an application-level coordination or idempotency strategy.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

A safer workflow to replay a Kafka message

  1. Identify the original. Record its source topic, partition, offset, key where appropriate, and failure metadata so the recovery action is tied to the right event.
  2. Fix or route around the cause. Correct malformed data or code, restore the unavailable dependency, or define a safe rejection path. Replaying before the cause is addressed can simply reproduce the failure.
  3. Choose a bounded replay set. Specify which records are eligible, how many will be replayed, and whether their order relative to newer records matters.
  4. Protect side effects. Make processing idempotent or deduplicate by a stable event identity; for Kafka-to-Kafka work, use transaction coordination where appropriate. For external effects, ensure the destination or application participates in the strategy.
  5. Observe recovery. Monitor the source consumer’s lag and the replay or DLQ flow, and check that recovered records do not create repeated failures or duplicate effects.

The replay shortcut worth fixing is rewinding a consumer or republishing a DLQ record without coordinating consumer position, output, ordering, and external side effects. Either action may be useful in a controlled recovery, but neither is a complete recovery plan by itself.

Product prices and availability are accurate as of the date/time indicated and are subject to change. Any price and availability information displayed on Amazon at the time of purchase will apply.