Streaming Kafka to Apache Iceberg
The streaming lakehouse is winning. Kafka topics land directly in Iceberg, and every analyst gets fresh, query-ready data without a nightly batch job. Here is what the modern pipeline looks like and what to watch.
Why Streaming to Iceberg Took Over
Iceberg won the open table format war. Snowflake, Databricks, Trino, DuckDB, Athena, BigQuery, Spark — they all read the same tables. The natural next question was how to get streaming data into Iceberg without a stitching layer of Spark jobs and Airflow DAGs. In 2026, the answer is: connect Kafka to Iceberg directly.
One Source of Truth
Streaming and batch consumers query the same tables. No more warehouse/lake divergence.
Minute-Level Freshness
Commits happen every 60 seconds by default. Analytics catch up to the operational store.
Open by Design
Iceberg is vendor-neutral. Your data survives platform changes and query-engine migrations.
The Three Delivery Paths
1. Kafka Connect Iceberg Sink
The community sink connector runs inside any Kafka Connect cluster. Configure schema handling, partitioning, and commit interval. Best when you already operate Connect and want direct control.
2. Confluent Tableflow
A managed materialization layer on Confluent Cloud that turns topics into Iceberg tables with no connector plumbing. Trades control for zero-ops.
3. WarpStream / AutoMQ Native Iceberg
Object-storage Kafka alternatives increasingly write Iceberg-compatible files as the storage format itself, eliminating the sink hop entirely.
Pipeline Metrics That Matter
Table Freshness Lag
Delta between latest Kafka offset and latest Iceberg snapshot timestamp. This is the number analysts feel.
Commit Interval Success Rate
Failed commits back up in-flight files. Alert when the last successful snapshot ages beyond 2x the configured interval.
Small File Count
Frequent commits create tiny files that wreck query performance. Track file count per partition and run compaction on a schedule.
Schema Drift Events
Iceberg supports evolution, but sinks fail on incompatible changes. Emit an alert when a topic schema shifts.
Design Guidelines
Partition by Ingest Time, Not Event Time
Prevents late-arriving records from rewriting historical partitions. Add event-time as a hidden partition transform for query pruning.
Tune Commit Interval to Consumer SLA
60 seconds is a good default. Sub-30s inflates small files; over 5 minutes hurts freshness.
Enforce Schema Compatibility
Schema Registry compatibility rules stop breaking changes at the producer. Iceberg readers depend on it.
Schedule Compaction and Snapshot Expiration
Iceberg tables need maintenance. Wire compaction and snapshot expiry into your orchestrator or use a managed catalog that does it for you.
Monitor Every Hop from Kafka to Iceberg
KLogic tracks connector health, commit lag, freshness, and schema drift across the full streaming lakehouse pipeline.