KRaft Migration: Monitoring Without ZooKeeper
Kafka 4.0 refuses to boot in ZooKeeper mode. Every remaining ZK-backed cluster has to migrate — and the metrics your dashboards depend on are changing underneath you. Here is what to watch during and after the cutover.
Why KRaft, and Why Now
KRaft replaces ZooKeeper with a Raft-based metadata log embedded in Kafka itself. It removes an entire distributed system from the operational surface, cuts metadata operations to milliseconds, and lets clusters scale to millions of partitions. With ZooKeeper mode removed in 4.0, migration stopped being optional.
One System, Not Two
No ZK ensemble to patch, back up, or size separately. Controllers colocate with (or run inside) brokers.
Faster Metadata
Controller failover drops from tens of seconds to under a second. Partition creation is orders of magnitude quicker.
Millions of Partitions
The old znode-per-partition ceiling is gone. Metadata replicates through the same log-based mechanism as your data.
Migration Playbook
1. Upgrade to a Migration-Capable Version
Get the cluster onto 3.7+ with ZK still active. Roll brokers first, then confirm inter-broker protocol is at the target version.
2. Provision the KRaft Controller Quorum
Start a 3-node controller quorum in migration mode. It reads existing metadata from ZooKeeper and begins replicating it into the KRaft log.
3. Enter Dual-Write Mode
Both ZK and KRaft receive writes. Monitor MigratingZkBrokerCount and metadata parity dashboards continuously.
4. Cut Over Brokers to KRaft
Roll each broker with zookeeper.metadata.migration.enable=false. Once the last broker flips, ZK is orphaned.
5. Decommission ZooKeeper
Wait one full retention window, verify no client is talking to ZK, then shut it down and remove from infra.
Metrics That Replace ZooKeeper Signals
ActiveControllerCount
Must be exactly 1 across the quorum. 0 or 2 signals an election in progress or a split.
MetadataLoaderLag
Bytes each broker lags behind the metadata log leader. Sustained lag delays partition-state changes.
CurrentControllerLeaderEpoch
Epoch changes indicate controller failovers. Frequent bumps mean quorum instability.
MigratingZkBrokerCount
Only relevant during migration. Should drift to zero as brokers finish cutover.
Alerts to Wire Before Cutover
Controller Quorum Loss
Alert if fewer than N/2+1 voters are healthy. This is the KRaft equivalent of losing the ZK ensemble.
Metadata Loader Lag Threshold
Any broker lagging > 10 MB or > 30 seconds should page. It cannot correctly apply new partition assignments.
Rapid Leader Elections
More than 3 leader changes per hour means the quorum is unstable — check GC, disk latency, and network to controller nodes.
Retire All ZooKeeper Alerts
Old ZK connection and znode watch alerts will go silent post-migration. Delete them explicitly or they become false negatives. See alerting best practices.
Post-Migration Checklist
Migrate to KRaft With Full Visibility
KLogic ships purpose-built KRaft dashboards and alerts so you can cut over from ZooKeeper without discovering blind spots the hard way.