KLogic
Architecture

Streaming Kafka to Apache Iceberg

The streaming lakehouse is winning. Kafka topics land directly in Iceberg, and every analyst gets fresh, query-ready data without a nightly batch job. Here is what the modern pipeline looks like and what to watch.

14 min read•September 10, 2026•KLogic Team

Why Streaming to Iceberg Took Over

Iceberg won the open table format war. Snowflake, Databricks, Trino, DuckDB, Athena, BigQuery, Spark — they all read the same tables. The natural next question was how to get streaming data into Iceberg without a stitching layer of Spark jobs and Airflow DAGs. In 2026, the answer is: connect Kafka to Iceberg directly.

One Source of Truth

Streaming and batch consumers query the same tables. No more warehouse/lake divergence.

Minute-Level Freshness

Commits happen every 60 seconds by default. Analytics catch up to the operational store.

Open by Design

Iceberg is vendor-neutral. Your data survives platform changes and query-engine migrations.

The Three Delivery Paths

1. Kafka Connect Iceberg Sink

The community sink connector runs inside any Kafka Connect cluster. Configure schema handling, partitioning, and commit interval. Best when you already operate Connect and want direct control.

2. Confluent Tableflow

A managed materialization layer on Confluent Cloud that turns topics into Iceberg tables with no connector plumbing. Trades control for zero-ops.

3. WarpStream / AutoMQ Native Iceberg

Object-storage Kafka alternatives increasingly write Iceberg-compatible files as the storage format itself, eliminating the sink hop entirely.

Pipeline Metrics That Matter

Table Freshness Lag

Delta between latest Kafka offset and latest Iceberg snapshot timestamp. This is the number analysts feel.

Commit Interval Success Rate

Failed commits back up in-flight files. Alert when the last successful snapshot ages beyond 2x the configured interval.

Small File Count

Frequent commits create tiny files that wreck query performance. Track file count per partition and run compaction on a schedule.

Schema Drift Events

Iceberg supports evolution, but sinks fail on incompatible changes. Emit an alert when a topic schema shifts.

Design Guidelines

Partition by Ingest Time, Not Event Time

Prevents late-arriving records from rewriting historical partitions. Add event-time as a hidden partition transform for query pruning.

Tune Commit Interval to Consumer SLA

60 seconds is a good default. Sub-30s inflates small files; over 5 minutes hurts freshness.

Enforce Schema Compatibility

Schema Registry compatibility rules stop breaking changes at the producer. Iceberg readers depend on it.

Schedule Compaction and Snapshot Expiration

Iceberg tables need maintenance. Wire compaction and snapshot expiry into your orchestrator or use a managed catalog that does it for you.

Monitor Every Hop from Kafka to Iceberg

KLogic tracks connector health, commit lag, freshness, and schema drift across the full streaming lakehouse pipeline.