KLogic
Architecture

Kafka Queues with KIP-932 Share Groups

KIP-932 finally gives Kafka point-to-point queue semantics. Multiple consumers can share a partition, acknowledge records individually, and retry failed ones — no SQS bridge required.

13 min read•September 21, 2026•KLogic Team

Consumer Groups vs Share Groups

Classic consumer groups pin one partition to one consumer at a time. Great for ordered stream processing, terrible for slow-record workloads where a single stuck message blocks everything behind it. Share groups flip the model: many consumers cooperate on the same partition, and each record is acknowledged individually.

Consumer Groups

  • • One partition per consumer
  • • Committed offsets advance monotonically
  • • Ordering guaranteed per partition
  • • Parallelism capped by partition count

Share Groups

  • • Many consumers per partition
  • • Per-record acknowledgement
  • • Ordering relaxed; retries out-of-order
  • • Parallelism scales with consumer count

When to Use Share Groups

Variable-Latency Workloads

Payment webhooks, external API calls, LLM inference — any workload where one slow record should not stall the whole partition.

Fine-Grained Retries

Individual records can be redelivered without rewinding an offset. Failed records go back into the queue with a delivery-attempt counter.

Elastic Consumers

Autoscale consumers past partition count. Perfect for spiky workloads that used to require a downstream queue.

Replacing SQS / RabbitMQ Bridges

Teams running Kafka plus a separate queue for work distribution can consolidate onto one platform.

How Share Groups Work

Delivery State Machine

Available — record fetched, waiting for a consumer
Acquired — consumer has it, ack timer running
Acknowledged — commit persisted, record retires
Released — nack or timeout, back to Available (attempt++)
Archived — delivery-attempt cap hit, poison pill

The broker tracks each record's state per share group. Configure share.record.lock.duration.ms for the ack timeout and share.delivery.count.limit for max attempts before archival.

Metrics You Must Monitor

Unacked Record Age

Oldest record in Acquired state per group. Rising values mean consumers cannot keep up or are hanging inside handlers.

Delivery Attempt Distribution

Percentiles of attempts per record. A p99 that trends toward the limit signals downstream instability.

Archived (Poison) Rate

Records that exhausted retries. Alert immediately — these bypass business logic and typically land in a DLQ topic.

In-Flight Records Per Consumer

Balance across consumers should stay even. Skew here indicates a slow handler or an autoscaler that under-provisioned.

Adoption Guardrails

1. Do Not Replace Streaming Consumers

Share groups relax ordering. Use them for work distribution, not for stateful stream processing where sequence matters.

2. Set Realistic Lock Durations

Lock timeouts shorter than handler p99 cause duplicate deliveries. Measure real latency before tuning.

3. Route Archives to a DLQ

Wire an archive handler that publishes to a dead-letter topic and pages on-call. Silent archives eat customer requests.

4. Alert on Both Lag and Attempt Count

Share-group lag alone hides retry storms. Combine with attempt-count percentiles using KLogic anomaly detection.

Monitor Share Groups With Confidence

KLogic tracks unacked-record age, delivery attempts, and archive rates across every share group so you can adopt Kafka queues without operating blind.