KLogic
Architecture

KRaft Migration: Monitoring Without ZooKeeper

Kafka 4.0 refuses to boot in ZooKeeper mode. Every remaining ZK-backed cluster has to migrate — and the metrics your dashboards depend on are changing underneath you. Here is what to watch during and after the cutover.

13 min read•September 5, 2026•KLogic Team

Why KRaft, and Why Now

KRaft replaces ZooKeeper with a Raft-based metadata log embedded in Kafka itself. It removes an entire distributed system from the operational surface, cuts metadata operations to milliseconds, and lets clusters scale to millions of partitions. With ZooKeeper mode removed in 4.0, migration stopped being optional.

One System, Not Two

No ZK ensemble to patch, back up, or size separately. Controllers colocate with (or run inside) brokers.

Faster Metadata

Controller failover drops from tens of seconds to under a second. Partition creation is orders of magnitude quicker.

Millions of Partitions

The old znode-per-partition ceiling is gone. Metadata replicates through the same log-based mechanism as your data.

Migration Playbook

1. Upgrade to a Migration-Capable Version

Get the cluster onto 3.7+ with ZK still active. Roll brokers first, then confirm inter-broker protocol is at the target version.

2. Provision the KRaft Controller Quorum

Start a 3-node controller quorum in migration mode. It reads existing metadata from ZooKeeper and begins replicating it into the KRaft log.

3. Enter Dual-Write Mode

Both ZK and KRaft receive writes. Monitor MigratingZkBrokerCount and metadata parity dashboards continuously.

4. Cut Over Brokers to KRaft

Roll each broker with zookeeper.metadata.migration.enable=false. Once the last broker flips, ZK is orphaned.

5. Decommission ZooKeeper

Wait one full retention window, verify no client is talking to ZK, then shut it down and remove from infra.

Metrics That Replace ZooKeeper Signals

ActiveControllerCount

Must be exactly 1 across the quorum. 0 or 2 signals an election in progress or a split.

MetadataLoaderLag

Bytes each broker lags behind the metadata log leader. Sustained lag delays partition-state changes.

CurrentControllerLeaderEpoch

Epoch changes indicate controller failovers. Frequent bumps mean quorum instability.

MigratingZkBrokerCount

Only relevant during migration. Should drift to zero as brokers finish cutover.

Alerts to Wire Before Cutover

Controller Quorum Loss

Alert if fewer than N/2+1 voters are healthy. This is the KRaft equivalent of losing the ZK ensemble.

Metadata Loader Lag Threshold

Any broker lagging > 10 MB or > 30 seconds should page. It cannot correctly apply new partition assignments.

Rapid Leader Elections

More than 3 leader changes per hour means the quorum is unstable — check GC, disk latency, and network to controller nodes.

Retire All ZooKeeper Alerts

Old ZK connection and znode watch alerts will go silent post-migration. Delete them explicitly or they become false negatives. See alerting best practices.

Post-Migration Checklist

All brokers report zookeeper.connected=false
Controller quorum shows 1 active leader
Metadata log fsync latency p99 under 10ms
Partition creation completes in under 1s
ZK ensemble stopped and snapshots archived
Runbooks updated to reference KRaft, not ZK

Migrate to KRaft With Full Visibility

KLogic ships purpose-built KRaft dashboards and alerts so you can cut over from ZooKeeper without discovering blind spots the hard way.