Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

Apache Kafka is a distributed event-streaming platform: producers publish records to named topics, Kafka stores those records according to retention settings, and consumers read them when needed. Topics are divided into partitions that can be distributed across brokers. This design supports parallel processing and replay, while ordering is guaranteed only within an individual partition—not across an entire topic.

What Kafka is—and what it is used for

Kafka is built to move and retain streams of events between systems. An event, also called a record or message, can contain a key, a value, a timestamp, and optional headers. A producer writes events to Kafka; a consumer reads them. Because those clients are decoupled, producers and consumers can operate and scale independently.

Kafka is useful when applications need to publish events for multiple consumers, process data continuously, or retain data so it can be read again. Unlike a one-read-and-remove model, Kafka keeps events until the topic’s retention policy removes them. A consumer can use its position in the stream to resume or replay data, subject to what remains available.

How Kafka’s main parts fit together

  1. Producers publish records to a topic. A record’s key can influence which partition receives it.
  2. Topics are named streams of records. Each topic can have many producers and many consumers.
  3. Partitions divide a topic into ordered sequences. Kafka preserves the write order within each topic-partition. Records with the same key are written to the same partition, which helps keep related records together.
  4. Brokers are the Kafka servers across which topic partitions are distributed. Partitions can have replicas on multiple brokers for fault tolerance.
  5. Consumers subscribe to topics and read records. A consumer’s offset represents its position in a partition, enabling it to continue from that position after a pause or restart.
  6. Consumer groups let consumers share work. Within one group, a partition is assigned to one consumer at a time; different partitions can be processed in parallel. Separate groups can read the same topic independently.

Ordering, scaling, and replay

Ordering is partition-scoped

Kafka provides a total order for each topic-partition, not a global order spanning every partition in a topic. If an application needs related events to be read in order, using the same key for those events commonly directs them to the same partition. That ordering guarantee does not extend to records placed in different partitions.

Free tools Windows power users keep installed

One-click scans. No signup required.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

Partitions set the parallelism limit

Partitions are the main unit of parallel work: producers can write to different partitions, and consumers in a group can process assigned partitions concurrently. A group cannot have more consumers actively assigned partitions than there are partitions to consume. Adding partitions may increase potential parallelism, but it changes the topic’s partitioning topology and can affect key distribution and assumptions about ordering.

Offsets make controlled rereading possible

An offset tracks a consumer’s position in a partition. Consumers can pause, restart, or replay from a chosen position while the relevant records remain within the topic’s retention window. Retention is therefore central to replay: Kafka does not preserve records indefinitely unless the topic’s retention settings and storage capacity support that outcome.

Replication and broker-failure resilience

Kafka replicates at the topic-partition level, keeping copies on multiple brokers. Replication can also span data centers or regions in supported deployments. The Kafka introduction describes a replication factor of three as a common production setting, meaning three copies of the data; it is an example, not a universal recommendation. The appropriate factor depends on the failures a deployment must tolerate, recovery needs, and storage cost.

Replication alone does not define the durability of every write. Producer acknowledgments and in-sync-replica settings also affect when a write is treated as successful and how the system behaves during failures. Those choices involve trade-offs between durability, availability, and latency, so the required behavior should be set according to the application’s failure and recovery requirements rather than assumed from the presence of replicas.

What’s actually slowing this PC down?

Pick the symptom - the matching free tool is one click away.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

Is Kafka a message queue, an event log, or a stream-processing platform?

Kafka can be described from more than one angle. Producers send messages and consumers receive them, so it can serve messaging use cases. Its retained, ordered partitions make it a distributed event log that consumers can reread. Kafka also provides APIs for stream processing, including Kafka Streams. Calling it only a conventional queue misses its retention-and-replay model; calling it only a processing engine misses the role of its stored event streams.

The distinction matters when choosing a system. Consider whether the application needs high throughput and partition-based parallelism, what ordering scope it requires, how long events must remain replayable, what replication and acknowledgment behavior it needs, and how much operational complexity it can support. Also consider how consumer groups and offsets will be managed, what delivery semantics are necessary, and whether integrated stream processing is useful.

Kafka APIs and stream processing

Kafka documents four API families: Admin, Producer, Consumer, and Kafka Streams. Admin supports management tasks; Producer and Consumer APIs publish and read records. Kafka Streams is a library for building stream-processing applications with transformations, joins, aggregations, windows, event-time processing, and stateful operations. Its processing model works with Kafka storage and offsets, which can support stronger processing guarantees than a loosely coupled external sink.

What Kafka’s delivery guarantees do—and do not—mean

Kafka applications can use at-most-once, at-least-once, or exactly-once patterns depending on producer, consumer, and processing configuration. The terms describe what can happen when processing and failures interact; they are not unconditional promises that every application receives exactly one copy of every event.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

Idempotent production and transactions

Kafka’s idempotent production mechanism uses producer IDs and sequence numbers. When a retry repeats a sequence number that is not the next expected number for that producer and topic-partition, the broker rejects it, preventing that retry from being written as a duplicate. Transactions can combine produced records and consumed offsets into an atomic unit.

The boundary of exactly-once processing

Kafka Streams can provide exactly-once processing. A transactional producer paired with a read-committed consumer can also support exactly-once processing across Kafka topics. These guarantees depend on the appropriate configuration and apply within the documented Kafka processing boundary. They do not automatically make an external database write, HTTP request, or other outside side effect exactly once; those effects need their own coordination or idempotency strategy.

Independent reader supportYour contribution helps us test, update, and keep practical guides available for everyone.Support on Ko-Fi

When Kafka is a good fit

  • Several applications need to consume the same retained event stream independently.
  • Consumers need the option to resume or replay events still covered by retention.
  • Work can be partitioned, and ordering requirements can be expressed within partitions, often by key.
  • The system needs configurable replication and delivery behavior, and the team can operate and monitor a distributed platform.
  • Stream transformations or stateful processing would benefit from Kafka Streams’ integrated model.

Kafka’s design is not a substitute for workload-specific decisions about retention, partitioning, replication, acknowledgments, transactions, and external side effects. Those settings determine whether its distributed log model matches the application’s ordering, recovery, and processing requirements.

Product prices and availability are accurate as of the date/time indicated and are subject to change. Any price and availability information displayed on Amazon at the time of purchase will apply.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.