Three technical shifts are shaping modern data engineering: Kafka is moving beyond ZooKeeper and evolving how consumers coordinate; Flink is developing cloud-oriented state management and higher-level SQL abstractions; and streaming ingestion is getting more closely connected to open lakehouse tables such as Iceberg. Together, they point toward a more integrated data stack—not a merger of the projects or proof that every organization is adopting it.
How Kafka, Flink, and Iceberg fit together
These projects address different parts of a data system. Kafka is commonly used to transport and retain event streams for consumers. Flink processes data, including stateful continuous workloads and batch jobs. Iceberg is an open table format that lets multiple compute engines work with data organized as tables. A system can use all three, but each still has its own role, operating requirements, and compatibility boundaries.
| Project | Primary role in this architecture | What is changing |
|---|---|---|
| Kafka | Event transport and consumer coordination | ZooKeeper-free KRaft operation and evolving consumer protocols |
| Flink | Stream and batch processing, including stateful computation | Cloud-oriented state, SQL and materialized-table features |
| Iceberg | Open table format for data in a lakehouse | Catalog-side planning and improvements to streaming scans |
The practical question is how these boundaries work together for a particular workload: how quickly data must be available, how it is recovered or replayed, and whether the connectors and versions in use support the required behavior.
Trend 1: Kafka is simplifying its operating model while expanding consumer options
KRaft makes Kafka independent of ZooKeeper
Apache Kafka 4.0, announced March 18, 2025, was the first major Kafka release designed to operate entirely without ZooKeeper, using KRaft by default. That is an architectural milestone: Kafka’s metadata and coordination no longer depend on a separate ZooKeeper ensemble. For operators, the change can simplify the system’s component layout, but it does not make a migration automatic. Existing clusters need a planned upgrade path, and teams should account for their deployment configuration and operational procedures.
#1 Best Overall
Kafka’s release index lists Kafka 4.2.2, dated September 29, 2026. That makes 4.0 useful for understanding the direction of the architecture, not a recommendation to select 4.0 as the current version. Check the release and upgrade documentation for the exact target version before changing a cluster.
Consumer coordination is becoming more flexible
Kafka 4.0 also made the next-generation consumer group protocol, KIP-848, generally available. The protocol is intended to improve rebalance behavior for consumer groups. The server enables the protocol, but clients opt in by setting group.protocol=consumer. Client support and configuration therefore matter: upgrading brokers alone does not mean all consumers have switched to the new protocol.
The same release introduced share groups as an early-access option for queue-like consumption. Because they were early access in Kafka 4.0, treat them as a feature to evaluate against the status and support of the Kafka version you plan to run, rather than assuming they are a settled replacement for existing consumption patterns.
What this means for a Kafka upgrade
- Plan the ZooKeeper-to-KRaft transition as an operational migration, not simply a broker-version bump.
- Check whether every relevant client supports the consumer protocol you intend to use, then opt in deliberately.
- Evaluate consumer behavior under your own workload; a protocol’s intended rebalance improvements are not a guarantee of a particular result for every group.
- Verify feature maturity and upgrade requirements against the exact Kafka release you will deploy.
Trend 2: Flink is making stateful processing more cloud-oriented and expressive
State management designed for cloud deployments
Apache Flink 2.0, announced March 24, 2025, introduced disaggregated state storage and management designed around remote distributed filesystems and cloud-native deployment constraints. The direction is to handle stateful processing in a way that better fits environments where compute and storage may be managed or scaled separately. It does not remove the need to understand a job’s state size, recovery requirements, or operational behavior.
Rank #2
SQL and materialized tables abstract more of the work
Flink 2.0 refined materialized tables so users can express business logic without directly managing the stream-versus-batch mechanics. It also optimized batch execution for workloads that do not need continuous processing. Those features support a broader choice of execution style; they do not mean every data pipeline should be continuously running or that SQL eliminates the design work around correctness and recovery.
Flink 2.3, announced June 25, 2026, extended the SQL capabilities with FROM_CHANGELOG and TO_CHANGELOG, and added more granular materialized-table evolution and refresh controls. It also included a native S3 filesystem marked experimental in that release. Teams considering that filesystem should treat its experimental status as a maturity caveat, not as a general production-readiness claim.
Upgrades still have compatibility costs
Flink 2.0 included breaking API and configuration changes, including removal of older APIs. An upgrade therefore involves more than comparing new features: review application code, configuration, connectors, and the migration path. Flink’s emphasis on cloud-oriented state and table abstractions is most relevant when it aligns with the workload and the team can absorb the associated upgrade and operational effort.
Trend 3: Streaming ingestion and open lakehouse tables are converging
Iceberg is extending table planning and streaming scans
Apache Iceberg 1.11.0, released May 19, 2026, advanced remote scan planning through the REST catalog. In the release description, server-side planning returns relevant scan tasks instead of requiring each client to fetch and inspect manifests. The work also extends to incremental Structured Streaming scans and metadata tables. These capabilities can change where planning work happens, but they do not by themselves establish a particular end-to-end freshness level or remove the need to operate a catalog and compatible compute clients.
Rank #3
Iceberg’s central role is different from Kafka’s and Flink’s: it is an open table format usable by multiple engines, rather than an event broker or a processing engine. Its evolution model also allows partition layouts to change over time: existing data retains its earlier partition specification while new data can use the new layout. Actual support and behavior depend on the engine and connector versions involved.
CDC and connector development are part of the bridge
Apache Flink CDC 3.6.0, announced March 30, 2026, supports Flink 1.20.x and 2.2.x. It adds an Oracle source and Hudi sink pipeline connectors, includes fixes across Iceberg and Kafka connectors, and advances schema-evolution work across several sources and sinks. These are signs of continued work on the plumbing needed to capture changes and connect systems; they are not evidence that every source, sink, or combination is production-ready for every use case.
Version compatibility needs particular attention. Iceberg 1.11.0 lists Flink 2.1 support and removes Flink 1.19 support, while Flink CDC 3.6.0 lists support for Flink 1.20.x and 2.2.x. Those statements describe different project components and do not establish that every combination works together. Confirm the precise connector, engine, catalog, and table-format versions for the intended pipeline before upgrading or deploying it.
Fluss is a related, distinct project
Apache Fluss graduated to an Apache top-level project on August 6, 2026. Its stated goal is unified streaming storage for real-time analytics in the lakehouse era. It is an example of the architectural direction toward bringing streaming storage closer to lakehouse use cases, but it is not another name for Kafka, Flink, or Iceberg, nor does its project status establish broad adoption.
Free tools Windows power users keep installed
One-click scans. No signup required.
Rank #4
How to evaluate a Kafka-to-Iceberg data path
A Kafka-to-Iceberg design typically involves more than moving records between two products. Consider the full path: how source changes are captured, what processes and transforms them, how the table is committed and discovered, and how downstream engines read updates. Flink and CDC connectors may be part of that path, but exact capabilities depend on the versions and connectors selected.
- Latency: Decide how fresh data needs to be and whether the workload needs continuous processing, scheduled refreshes, or batch execution.
- State and recovery: Establish what the processing job must retain, how it recovers, and whether replaying source events is part of the recovery plan.
- CDC and schemas: Verify support for the source database, change types, schema changes, and destination behavior in the exact connector versions.
- Tables and partitions: Check how the engine handles table commits, metadata, and partition evolution, especially if more than one engine reads or writes the table.
- Compatibility: Validate the Kafka, Flink, CDC, Iceberg, and catalog versions as a complete combination; a feature listed by one project does not guarantee support across the stack.
- Operational burden: Compare the effort of managing brokers, processing jobs, catalogs, connectors, and copies of data with the freshness and replay requirements the architecture serves.
There is no single performance ranking in these release announcements that can settle those choices. A meaningful comparison needs a workload-matched evaluation, including data volume, latency target, state size, failure and recovery conditions, and the actual software versions.
What these trends do—and do not—show
The releases show technical direction: Kafka is advancing a ZooKeeper-free operating model and consumer coordination; Flink is developing cloud-oriented state and higher-level processing abstractions; and Iceberg, CDC, and connector projects are improving the paths between streams and tables. They do not show that the tools have merged, that lakehouse data is universally real-time or zero-copy, or that these designs have become the default across the industry. Treat the developments as capabilities to assess against specific workload needs, not as proof of market-wide adoption.
Quick Recap
Product prices and availability are accurate as of the date/time indicated and are subject to change. Any price and availability information displayed on Amazon at the time of purchase will apply.
Recommended Free Tools

