Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

A Kafka consumer can process the same record again after a restart or partition reassignment if the business effect succeeds but the offset commit does not. That is expected at-least-once behavior—not a promise that every event appears exactly twice. The durable fix is to make repeated work harmless or to commit the result and progress within one transaction boundary.

Why a Kafka consumer replays a record

Kafka tracks consumer progress with offsets. A committed offset identifies the next record the consumer should read, not the last one it completed. Processing a record and committing the offset are separate actions unless the application coordinates them atomically.

  1. The consumer polls a record at offset N.
  2. The application applies a database change or another external side effect.
  3. The consumer commits offset N+1, the next position to read.

If the process stops after step 2 but before step 3, Kafka still has the earlier committed position. After restart or reassignment, the new consumer reads from that position and can apply the effect for N again. The same issue can affect a batch if the application commits progress only after processing the batch. Apache Kafka describes at-least-once delivery as the default: “Otherwise, Kafka guarantees at-least-once delivery by default, and allows the user to implement at-most-once delivery by disabling retries on the producer and committing offsets in the consumer prior to processing a batch of messages.” (Apache Kafka 4.1 Design documentation)

Committing before doing the work reverses the risk: if the process stops after the commit but before the side effect, the record may not be read again, so the work is lost. For systems where losing an event is unacceptable, the usual trade-off is to commit after successful work and make replays safe.

What’s actually slowing this PC down?

Pick the symptom - the matching free tool is one click away.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

Choose a strategy based on the output system

The right approach depends on what the consumer writes to and on the business cost of loss versus replay. These strategies have different guarantee boundaries:

Strategy Loss versus replay Destination participation Guarantee scope Operational considerations
Idempotent business operation Replay is safe if repeating the operation yields the same final state. Works with an external database or API if it supports an appropriate stable key or idempotency mechanism. The repeated business operation, not Kafka delivery itself. Requires a suitable key and careful handling of non-idempotent actions.
Manual commit after success Failure before commit can replay completed work; committing only after success avoids skipping unfinished work. The sink need not join a Kafka transaction, but it must tolerate a replay or coordinate with progress storage. At-least-once processing with offsets controlled by the application. Application must commit the correct next offset and manage partition progress.
Atomic sink-side result and offset Result and progress succeed or fail together within the sink transaction. The destination must be able to store both the business result and Kafka progress atomically. The transaction boundary of that destination. Requires sink-specific offset storage and recovery logic.
Kafka transactions or Kafka Streams Kafka output and consumed offsets can be committed together, avoiding a visible partial Kafka-to-Kafka result. Kafka output topics and, with Streams, its managed state stores. Kafka processing state and output covered by the transaction; not arbitrary external calls. Requires transactional configuration and correct handling of transaction aborts and consumer position.
Commit before processing A crash after commit but before the side effect can lose work; it avoids replay of records already committed. No sink transaction required. At-most-once handling of committed records. Usually unsuitable when event loss is unacceptable.

Make external business effects idempotent

For a database or HTTP destination, use a stable event ID or business key so a retry or Kafka replay does not create a second business effect. Kafka’s design guide gives primary-key overwrites as an example of an idempotent update. An upsert keyed by an event identifier can have that property; a blind counter increment, payment, or append-only insert usually does not.

When the destination is a relational database, one common implementation is to store a processed-event marker and apply the business mutation in the same database transaction. For example, insert the event ID into a table with a unique constraint, then perform the business update only if the insert was new; commit both changes together. This schema pattern is implementation guidance, not a Kafka-prescribed design. If the transaction rolls back, neither the marker nor the mutation remains; if it commits, a replay encounters the existing ID and can be treated as already handled.

For an external API, use its idempotency-key facility when available, using the stable event or operation ID. If the endpoint has no such mechanism and the action cannot be safely repeated, Kafka offsets alone cannot make that external effect exactly once; the application needs a destination-side coordination design appropriate to that API.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

Commit offsets only as far as completed work

With manual commits, wait for successful processing before advancing progress. KafkaConsumer documentation specifies that when committing a particular record manually, the committed offset should be that record’s offset plus one, because the committed value is the next record to read. Its example describes a new consumer repeating the last batch’s insert when failure occurs before the commit. (Apache KafkaConsumer 2.8.1 API documentation)

Be especially careful with parallel work within a partition. If offset 12 is still processing, do not commit offset 13 just because offset 13 or a later record finished; doing so could skip unfinished work after a crash. Track completed offsets per partition and advance the committed position only through the highest contiguous range that is finished.

If processing is moved off the thread that calls poll, the consumer still needs to poll within max.poll.interval.ms to remain live. Keep polling and coordinating completed work rather than committing a position that has advanced beyond the work actually finished. The exact APIs and configuration details should match the Kafka client version deployed; the cited consumer guidance is for version 2.8.1.

Understand what Kafka’s exactly-once features cover

Producer idempotence handles a different duplicate problem

Producer idempotence prevents duplicates caused by producer retries within a producer session; it does not deduplicate consumer-side database effects or arbitrary application resends. KafkaProducer 3.9.2 documents that, starting with Kafka 3.0, enable.idempotence defaults to true in that client API, with related retry and acknowledgment defaults configured accordingly. These are version-specific Java client details, not a guarantee about every client language or older release. (KafkaProducer 3.9.2 API documentation)

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

Kafka transactions coordinate Kafka input and output

For Kafka-to-Kafka processing, a transactional producer can write output records and commit the consumed input offsets in the same Kafka transaction. The current Kafka 4.1 design guide’s workflow includes a transactional.id, disabling auto-commit with enable.auto.commit=false, and configuring downstream consumers with isolation.level=read_committed. A consumer that does not use read-committed isolation may see records from aborted or still-open transactions. (Apache Kafka 4.1 Design documentation)

This guarantee covers the Kafka records and offsets in that transaction. It does not roll back or atomically commit an unrelated database write or external API request. An external destination needs to cooperate—for example, by atomically storing its result and the consumed offset in its own transaction.

Kafka Streams integrates Kafka state and output

Kafka Streams integrates processing with its state stores, output topics, and offsets, so these can be handled atomically within its Kafka processing model. The cited Kafka Streams 2.1 guide describes this behavior and calls the question of processing each record once despite failures a frequently asked question. Its configuration names and release details are old; check the documentation for the deployed Kafka Streams version before applying configuration. (Apache Kafka Streams 2.1 Core Concepts)

Independent reader supportYour contribution helps us test, update, and keep practical guides available for everyone.Support on Ko-Fi

Is auto-commit the cause?

Not by itself. Automatic commits can provide at-least-once behavior if the application finishes processing every record returned by poll before calling poll again or closing the consumer. If the application allows an automatic commit to get ahead of unfinished work, a crash can cause records to be skipped. The important question is whether committed progress can advance beyond completed processing, not simply whether auto-commit is enabled. Consult the API documentation for the exact client version in use.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

A practical decision path

  1. Identify the side effect. Determine whether the consumer writes to Kafka, a database, an API, or more than one destination.
  2. Decide which failure is less acceptable. If losing an event is unacceptable, do not commit its offset before successful work; design for replay.
  3. For an external sink, establish replay safety. Prefer an idempotent operation keyed by a stable event or business ID. Where needed, commit a deduplication marker and the business mutation in the same sink transaction.
  4. For Kafka output, use Kafka’s transaction boundary. Coordinate output records with input offsets, disable auto-commit for the transactional flow, and have downstream consumers use read-committed isolation.
  5. Check offset advancement. Commit the next offset only after the corresponding work is complete, and in parallel processing do not pass unfinished offsets in a partition.

Kafka’s design documentation and the KafkaConsumer API documentation explain the offset and delivery model. For a broader treatment, O’Reilly’s Kafka: The Definitive Guide, 2nd Edition includes consumer offsets, reliable delivery, idempotent producers, and exactly-once semantics.

Product prices and availability are accurate as of the date/time indicated and are subject to change. Any price and availability information displayed on Amazon at the time of purchase will apply.