Apache Kafka 4.0: New Features and Monitoring Impact
Kafka 4.0 is the biggest release in years. ZooKeeper is finally gone, the consumer protocol has been rewritten, and share groups add true queue semantics. Here is what changed and what your monitoring stack needs to know.
Why Kafka 4.0 Is a Turning Point
Kafka has been iterating steadily since 3.0, but 4.0 removes long-deprecated components and turns several preview features into defaults. That means clusters upgraded from 3.x pick up new metrics, new failure modes, and new operational patterns whether teams plan for them or not.
KRaft is Default
ZooKeeper mode is removed. New clusters run KRaft, and 3.x clusters must migrate before upgrading.
New Consumer Protocol
KIP-848 replaces the old rebalance protocol with an incremental, server-driven model.
Share Groups (Queues)
KIP-932 delivers point-to-point queue semantics with cooperative offsets and per-record acknowledgement.
Feature Deep Dive
KRaft as the Only Metadata Layer
ZooKeeper support is gone. Controller quorum health, active-controller elections, and metadata log lag are now core signals. See our KRaft migration guide for the exact metrics to alert on.
KIP-848 Next-Gen Consumer Rebalancing
The broker now owns rebalance decisions and streams assignment deltas to clients. Long stop-the-world pauses shrink to sub-second incremental moves, but the metrics you graphed for years no longer mean the same thing.
KIP-932 Share Groups
Multiple consumers can now cooperatively process a single partition with per-record ack semantics. Kafka finally competes head-on with SQS and RabbitMQ for queue workloads without extra middleware.
Transactions v2
A rewritten transaction coordinator reduces producer-side epoch bumps and cuts end-to-end latency for exactly-once workloads. New JMX metrics expose per-transaction commit times.
Tiered Storage GA at Scale
Tiered storage graduated from preview and is production-hardened. Learn the operational patterns in our tiered storage monitoring guide.
What Your Monitoring Stack Needs to Change
Every upgrade to 4.0 introduces new JMX beans, retires old ones, and reshapes familiar dashboards. Miss the transition and blind spots appear in the exact places you rely on for capacity planning.
Add Controller Quorum Metrics
Track MetadataLoaderLag, ActiveControllerCount, and per-voter fetch lag. These replace ZK session and znode watchers.
Reinterpret Rebalance Metrics
Under KIP-848, RebalanceRate stays flat while assignment churn shows up in new server-side counters. Update thresholds accordingly.
Track Share Group Lag Separately
Share groups do not use committed offsets the same way. Alert on unacked-record age and delivery attempts, not partition lag alone.
Watch Remote Storage IO
Tiered storage adds RemoteBytesInPerSec and remote fetch latency. Budget alerts for both hot and archive tiers.
Upgrade Checklist
1. Complete KRaft Migration First
4.0 refuses to boot in ZooKeeper mode. Finish your KRaft migration on 3.7+ before attempting the version bump.
2. Roll Clients Ahead of Brokers
The new consumer protocol is opt-in per group. Upgrade client libraries, then flip group.protocol=consumer per workload.
3. Baseline New Metrics
Capture 7 days of pre-upgrade metrics for consumer lag, request latency, and controller health so post-upgrade regressions stand out.
4. Rewrite Old Alerts
Retire ZooKeeper-based alerts and rebalance-count alerts. Replace with controller quorum health and KIP-848 assignment-churn signals via KLogic alerting.
Stay Ahead of Kafka 4.0
KLogic ships day-one dashboards for KRaft, KIP-848 consumers, share groups, and tiered storage. Upgrade with confidence.