KLogic
Performance

Kafka Tiered Storage: Setup, Monitoring, and Cost Savings

Broker EBS is expensive, and 90% of Kafka data is cold. Tiered storage offloads old segments to object storage — done right, it cuts infra bills by more than half while keeping historical replay a single fetch away.

15 min read•September 15, 2026•KLogic Team

Why Tiered Storage Now

KIP-405 landed in preview back in Kafka 3.6 and is fully production-ready in 4.x. Tiered storage lets brokers keep only a hot window on local disk and delegate everything older to S3, GCS, or Azure Blob. The savings compound at scale — a 200 TB cluster typically costs 60–80% less to run.

Smaller Brokers

Local disks size for hot data only. Broker footprint shrinks and rebalances get dramatically faster.

Object Storage Pricing

S3 is roughly 10x cheaper per GB than EBS gp3. For long retention, that difference dwarfs everything else.

Infinite Retention

Weeks of retention become months without buying new brokers. Compliance replay stops being a capacity conversation.

Enabling Tiered Storage

Broker Configuration

remote.log.storage.system.enable=true
remote.log.storage.manager.class.name=org.apache.kafka.server.log.remote.storage.RemoteLogManager
remote.log.metadata.manager.listener.name=PLAINTEXT
rsm.config.storage.bucket.name=kafka-tiered-prod
rsm.config.storage.aws.region=us-east-1

Topic Configuration

remote.storage.enable=true
local.retention.ms=86400000 # 1 day hot
retention.ms=2592000000 # 30 days total
segment.bytes=536870912 # 512 MB

Metrics That Only Exist with Tiered Storage

RemoteCopyBytesPerSec

Rate at which segments upload to object storage. Sudden drops signal a stuck uploader or IAM regression.

RemoteFetchLatencyAvg

Time to serve a historical read from tiered storage. Alert on p95 above the SLO for compliance replay.

RemoteLogSizeBytes

Total bytes in remote tier per topic-partition. Feeds your object-storage bill forecast.

RemoteCopyLagBytes

Local segments waiting for upload. Sustained lag means local disks will fill even though remote is enabled.

Operational Patterns That Actually Work

Right-Size Local Retention

Match local.retention.ms to consumer lag p99 plus a safety buffer. Too short and every restart hits object storage.

Pre-Warm After Failover

A newly-leader broker has no local hot data. Trigger a background prefetch so first-consumer reads do not stampede S3.

Budget Object-Storage Request Cost

S3 GETs add up. Coalesce fetches with larger segments (512 MB+) and monitor request rate alongside byte volume.

Correlate with Cost Data

Link RemoteLogSizeBytes to your cloud bill through cost optimization dashboards so retention decisions have a dollar sign.

Common Failure Modes

IAM changes silently break uploads
Local disk fills because upload lag went unnoticed
Historical replay saturates a single broker
S3 throttling under bulk-consumer workloads
Cost surprise from tiny segment sizes
Region mismatch between broker and bucket

Turn Tiered Storage Into Real Savings

KLogic surfaces upload lag, remote-fetch latency, and per-topic remote size so you can adopt tiered storage without operational blind spots.