Kafka Queues with KIP-932 Share Groups
KIP-932 finally gives Kafka point-to-point queue semantics. Multiple consumers can share a partition, acknowledge records individually, and retry failed ones — no SQS bridge required.
Consumer Groups vs Share Groups
Classic consumer groups pin one partition to one consumer at a time. Great for ordered stream processing, terrible for slow-record workloads where a single stuck message blocks everything behind it. Share groups flip the model: many consumers cooperate on the same partition, and each record is acknowledged individually.
Consumer Groups
- • One partition per consumer
- • Committed offsets advance monotonically
- • Ordering guaranteed per partition
- • Parallelism capped by partition count
Share Groups
- • Many consumers per partition
- • Per-record acknowledgement
- • Ordering relaxed; retries out-of-order
- • Parallelism scales with consumer count
When to Use Share Groups
Variable-Latency Workloads
Payment webhooks, external API calls, LLM inference — any workload where one slow record should not stall the whole partition.
Fine-Grained Retries
Individual records can be redelivered without rewinding an offset. Failed records go back into the queue with a delivery-attempt counter.
Elastic Consumers
Autoscale consumers past partition count. Perfect for spiky workloads that used to require a downstream queue.
Replacing SQS / RabbitMQ Bridges
Teams running Kafka plus a separate queue for work distribution can consolidate onto one platform.
How Share Groups Work
Delivery State Machine
The broker tracks each record's state per share group. Configure share.record.lock.duration.ms for the ack timeout and share.delivery.count.limit for max attempts before archival.
Metrics You Must Monitor
Unacked Record Age
Oldest record in Acquired state per group. Rising values mean consumers cannot keep up or are hanging inside handlers.
Delivery Attempt Distribution
Percentiles of attempts per record. A p99 that trends toward the limit signals downstream instability.
Archived (Poison) Rate
Records that exhausted retries. Alert immediately — these bypass business logic and typically land in a DLQ topic.
In-Flight Records Per Consumer
Balance across consumers should stay even. Skew here indicates a slow handler or an autoscaler that under-provisioned.
Adoption Guardrails
1. Do Not Replace Streaming Consumers
Share groups relax ordering. Use them for work distribution, not for stateful stream processing where sequence matters.
2. Set Realistic Lock Durations
Lock timeouts shorter than handler p99 cause duplicate deliveries. Measure real latency before tuning.
3. Route Archives to a DLQ
Wire an archive handler that publishes to a dead-letter topic and pages on-call. Silent archives eat customer requests.
4. Alert on Both Lag and Attempt Count
Share-group lag alone hides retry storms. Combine with attempt-count percentiles using KLogic anomaly detection.
Monitor Share Groups With Confidence
KLogic tracks unacked-record age, delivery attempts, and archive rates across every share group so you can adopt Kafka queues without operating blind.