Table of Contents

    consumer groups

    MESSAGING & ASYNCHRONOUS PROCESSING

    Consumer Groups

    Learn how Kafka consumer groups allow several consumers to process topic partitions in parallel, how group IDs create independent subscriptions, how partitions are assigned and rebalanced, how offsets track progress, and how lag, duplicate processing, failures, scaling, hot partitions, poison events, retries, static membership, idempotency, and observability affect production consumers.

    Introduction

    A Kafka topic can contain a continuous stream of events. One consumer can process that stream, but a single consumer eventually reaches limits in CPU, memory, network, database access, or event-processing capacity.

    A consumer group allows several consumer instances belonging to the same application to cooperate.

    Kafka Topic
        |
        +-- Partition 0
        +-- Partition 1
        +-- Partition 2
        +-- Partition 3
                  |
                  v
            Consumer Group
                  |
                  +-- Consumer A
                  +-- Consumer B

    Kafka assigns topic partitions among the active members of the group. Each partition is processed by one active consumer within that group at a time.

    Core idea: Consumers in the same group divide partition processing. Consumers in different groups independently read the same retained topic records.

    Prerequisites

    # Prerequisite Why It Is Needed
    1 Kafka topics and partitions Partitions are the units distributed among consumers in a group.
    2 Kafka producers and consumers Producers append records, while consumers retrieve and process them.
    3 Offsets A consumer group tracks its progress separately for every partition.
    4 Idempotency Failures and offset recovery can cause an event to be processed again.
    5 Partition ordering Kafka ordering is scoped to an individual partition.
    6 Backpressure Consumer groups can fall behind when event production exceeds processing capacity.
    7 Observability Consumer lag, processing failures, offset commits, and rebalances require monitoring.

    What Is a Consumer Group?

    A Kafka consumer group is a collection of consumers that work together to process records from one or more topics.

    Every consumer in the group uses the same group identifier, commonly configured through group.id.

    group.id = certificate-processors
    
    
    Members:
    
    certificate-worker-1
    
    certificate-worker-2
    
    certificate-worker-3

    Kafka assigns partitions among the group members so that the members can process different partitions concurrently without intentionally repeating the same partition work inside the group.

    Consumer-group Flow
    consumer joins group → group discovers members → partitions are assigned → consumers process records → offsets record progress

    Group ID

    The group ID identifies the logical consumer application.

    Topic:
    
    learning-domain-events
    
    
    Consumer Group:
    
    notification-service
    
    
    Consumers:
    
    notification-1
    notification-2
    notification-3

    Consumers using the same group ID cooperate and divide available partitions.

    Consumers using different group IDs represent independent subscriptions.

    Topic:
    
    learning-domain-events
    
    
    Group 1:
    
    notification-service
    
    
    Group 2:
    
    analytics-service
    
    
    Group 3:
    
    search-index-service

    Each group can process the complete topic independently and maintain its own offsets.

    Group-ID rule: Use the same group ID when application instances should share processing. Use different group IDs when applications need independent copies of the event stream.

    Same Group vs Different Groups

    Configuration Behaviour
    Same group ID Consumers cooperate and divide partition ownership.
    Different group IDs Consumers read the topic independently and maintain separate offsets.
    One consumer, one group The consumer processes all assigned topic partitions.
    Several consumers, one group Partitions are distributed among active group members.
    Several consumer groups Every group can independently process the retained stream.

    Partitions Determine Parallelism

    Kafka distributes partitions, not individual records, among active members of a consumer group.

    Four Partitions and One Consumer

    Partition 0 --+
    Partition 1 --+
    Partition 2 --+--> Consumer A
    Partition 3 --+

    One consumer processes all four partitions.

    Four Partitions and Two Consumers

    Partition 0 --+
                  +--> Consumer A
    Partition 1 --+
    
    
    Partition 2 --+
                  +--> Consumer B
    Partition 3 --+

    Four Partitions and Four Consumers

    Partition 0 -> Consumer A
    
    Partition 1 -> Consumer B
    
    Partition 2 -> Consumer C
    
    Partition 3 -> Consumer D

    Four Partitions and Six Consumers

    Partition 0 -> Consumer A
    
    Partition 1 -> Consumer B
    
    Partition 2 -> Consumer C
    
    Partition 3 -> Consumer D
    
    
    Consumer E -> No partition
    
    Consumer F -> No partition

    Consumers without partition assignments remain inactive for ordinary record processing in that group.

    Maximum Active Consumers

    For a group consuming one topic, a simplified upper bound for active partition-processing consumers is:

    \[ MaximumActiveConsumers = NumberOfTopicPartitions \]

    When a group subscribes to several topics, applicable partition assignments depend on the combined subscribed partitions and the selected assignment strategy.

    Adding consumers beyond available partition assignments improves neither active parallelism nor throughput.

    Ordering within a Consumer Group

    Kafka preserves the order of records within one partition.

    Partition for learner-1042:
    
    Offset 100 -> AccountRegistered
    
    Offset 101 -> CourseEnrolled
    
    Offset 102 -> LessonCompleted
    
    Offset 103 -> CourseCompleted

    The consumer assigned to this partition reads the records in partition order.

    Records in separate partitions do not have one automatic global order.

    Partition 0:
    
    Event A1
    Event A2
    Event A3
    
    
    Partition 1:
    
    Event B1
    Event B2
    Event B3
    
    
    No automatic total order exists
    between Event A2 and Event B2.

    Partition Key and Group Processing

    A producer commonly uses a business key to place related records in the same partition.

    Kafka record key:
    
    learner-1042
    
    
    Related events:
    
    LearnerRegistered
    LearnerEnrolled
    LessonCompleted
    CourseCompleted
    
    
    Result:
    
    Events can be routed to
    one partition and processed
    in partition order.

    A poor key can create a hot partition, while a random key can break required per-entity ordering.

    Consumer Offsets

    An offset identifies a record's position inside a partition.

    Partition 2:
    
    Offset 8,000 -> Event A
    
    Offset 8,001 -> Event B
    
    Offset 8,002 -> Event C

    Each consumer group maintains its own offset progress for each partition.

    Topic Partition 2:
    
    
    Notification Group:
    
    Committed offset = 8,002
    
    
    Analytics Group:
    
    Committed offset = 7,950
    
    
    Search Group:
    
    Committed offset = 7,700

    One slow consumer group does not change the positions of the other groups.

    Offset Commit

    An offset commit records the consumer group's processing position.

    Consumer fetches:
    
    Offsets 100 to 109
    
    
    Consumer processes:
    
    Offsets 100 to 109
    
    
    Consumer commits:
    
    Next required offset = 110

    After restarting or receiving the partition again, a consumer can resume from the group's last committed position.

    Auto Commit vs Manual Commit

    Mode General Behaviour Main Risk
    Automatic offset commit The client commits offsets according to configured automatic behaviour. Progress can be committed before the complete business effect is durable if processing is not aligned correctly.
    Manual synchronous commit The consumer requests an offset commit and waits for its result. Increases commit-path latency.
    Manual asynchronous commit The consumer requests a commit without blocking for the complete response path. Commit failures and callback ordering require careful handling.

    Commit rule: Commit progress only after the application reaches the intended durable processing boundary.

    Commit before Processing

    Consumer fetches event
          |
          v
    Consumer commits offset
          |
          v
    Consumer begins processing
          |
          v
    Consumer crashes
          |
          v
    Restart begins after event
          |
          v
    Business operation is skipped.

    Committing too early can produce an at-most-once style failure in which required work is not completed.

    Commit after Processing

    Consumer fetches event
          |
          v
    Consumer processes event
          |
          v
    Business transaction commits
          |
          v
    Consumer crashes before
    offset commit
          |
          v
    Event is delivered again.

    Committing after processing protects against silently skipped work but can produce duplicate processing. The consumer must be idempotent.

    Idempotent Consumer

    An idempotent consumer produces one intended business outcome even when the same event is delivered more than once.

    {
      "eventId": "stable-unique-event-id",
      "eventType": "LearnerEnrolled",
      "eventVersion": 1,
      "tenantId": 17,
      "learnerId": 1042,
      "courseId": 42
    }

    The event identifier can be recorded in the same transaction as the consumer's business update.

    Deduplication Table

    CREATE TABLE processed_consumer_events
    (
        consumer_group VARCHAR(150) NOT NULL,
        event_id VARCHAR(150) NOT NULL,
        processed_at TIMESTAMP NOT NULL,
    
        PRIMARY KEY
        (
            consumer_group,
            event_id
        )
    );

    The unique constraint prevents the same consumer group from applying the same event twice when the check and business operation use an appropriate transactional design.

    Rebalancing

    Rebalancing redistributes partition assignments among active members of a consumer group.

    A rebalance can occur when:

    • A consumer joins the group
    • A consumer leaves the group
    • A consumer fails or becomes unresponsive
    • The subscribed topic's partition count changes
    • The group's topic subscriptions change
    • Group membership or assignment configuration changes
    Before:
    
    Consumer A -> Partitions 0 and 1
    
    Consumer B -> Partitions 2 and 3
    
    
    Consumer C joins
          |
          v
    Group rebalances
          |
          v
    
    Possible result:
    
    Consumer A -> Partitions 0 and 1
    
    Consumer B -> Partition 2
    
    Consumer C -> Partition 3

    Consumer Failure

    Consumer B owns:
    
    Partitions 2 and 3
    
    
    Consumer B stops responding
          |
          v
    Group membership detects failure
          |
          v
    Partitions are reassigned
          |
          v
    Another consumer resumes
    from committed offsets.

    Records processed after the last committed offset can be processed again by the replacement consumer.

    Heartbeats and Session Activity

    Consumers use group-management communication to keep their membership active. Heartbeats help the group determine whether members are still participating.

    Consumer
        |
        +-- Fetch records
        +-- Process records
        +-- Maintain group activity
        +-- Commit offsets

    If the group concludes that a member is unavailable, the member's partitions can be reassigned.

    Slow Processing and Polling

    Long processing operations can interfere with healthy group participation when the consumer does not interact with Kafka within its configured processing boundaries.

    Consumer polls records
          |
          v
    One record takes a long time
          |
          v
    Consumer does not poll again
    within expected boundary
          |
          v
    Group can consider the member
    unable to continue normally
          |
          v
    Rebalance can occur.

    Possible design directions include:

    • Reduce records returned per poll
    • Bound event-processing duration
    • Move long-running work to a separate task queue
    • Pause assigned partitions while bounded work completes
    • Adjust consumer configuration using measured processing behaviour

    Rebalance Impact

    During partition reassignment, consumers must stop processing revoked partitions and begin processing newly assigned partitions from valid positions.

    Frequent rebalances can cause:

    • Temporary processing interruptions
    • Repeated initialization work
    • Cache rebuilding inside consumers
    • Duplicate processing around offset boundaries
    • Increased consumer lag
    • Reduced overall throughput

    Partition-assignment Strategies

    A partition-assignment strategy determines how partitions are distributed among group members.

    Strategy Direction General Goal
    Range-style assignment Assign contiguous topic-partition ranges to consumers.
    Round-robin assignment Distribute subscribed partitions across consumers in rotation.
    Sticky assignment Balance partitions while trying to preserve prior assignments.
    Cooperative sticky assignment Move affected assignments incrementally rather than revoking all ownership at once where supported.

    Final selection depends on subscription patterns, Kafka client support, deployment behaviour, balance requirements, and the cost of partition movement.

    Static Membership

    Dynamic consumers can receive new generated member identities when they restart. Static membership allows stable consumer identities to be associated with application instances where supported and configured.

    Stable membership can help reduce unnecessary partition movement during controlled short restarts, but failed or duplicated member identities still require correct lifecycle handling.

    Consumer instance:
    
    group.id = search-indexers
    
    group.instance.id = search-indexer-03

    Consumer Lag

    Consumer lag measures how far a consumer group remains behind the latest available partition position.

    A simplified record-based calculation is:

    \[ ConsumerLag = LogEndOffset - CurrentCommittedOffset \]

    Partition log-end offset:
    
    50,000
    
    
    Consumer-group offset:
    
    47,500
    
    
    Consumer lag:
    
    2,500 records

    Lag should be measured per partition. One heavily loaded partition can fall behind even when the group's total or average lag appears acceptable.

    Event-age Lag

    Record count does not directly express business delay because events can differ in size and processing cost.

    Group A:
    
    10,000 small events behind
    
    
    Group B:
    
    100 expensive events behind

    Monitor the age of the oldest unprocessed event as well as offset distance.

    Lag Growth

    Let \(P\) represent the event-production rate and \(C\) represent the consumer group's successful processing rate.

    Lag grows when:

    \[ P > C \]

    A simplified lag-growth rate is:

    \[ LagGrowthRate = P - C \]

    A group can drain its backlog only when its sustained completion rate exceeds the incoming production rate for the affected partitions.

    Estimated Drain Time

    When consumer capacity exceeds incoming production, a simplified estimate is:

    \[ EstimatedDrainTime = \frac{ CurrentLag }{ ConsumerRate - ProducerRate } \]

    This is a planning approximation. Event cost, retries, partition skew, consumer rebalances, and downstream dependencies affect actual recovery.

    Hot Partitions

    Adding group members cannot divide one Kafka partition among several active consumers in the same group.

    Topic has four partitions:
    
    
    Partition 0:
    
    5,000 events per second
    
    
    Partition 1:
    
    100 events per second
    
    
    Partition 2:
    
    120 events per second
    
    
    Partition 3:
    
    90 events per second

    The consumer assigned to Partition 0 becomes the group's bottleneck. Consumers assigned to other partitions can remain underused.

    Possible controls include:

    • Choose a more evenly distributed producer key
    • Split safely independent work across controlled subkeys
    • Increase partitions before current capacity is exhausted
    • Separate the heavy workload into another topic
    • Optimize processing for the hot event class
    • Apply producer-side rate or admission limits

    Scaling rule: Adding consumers helps only when Kafka can assign additional partitions containing useful work to those consumers. It does not divide one hot partition automatically.

    Poison Events

    A poison event repeatedly fails because of invalid data, unsupported schema, a code defect, or an unavailable dependency.

    Consumer reads Offset 500
          |
          v
    Processing fails
          |
          v
    Consumer retries Offset 500
          |
          v
    Processing fails again
          |
          v
    Later partition records
    cannot progress safely.

    Use bounded retries and a controlled failure path. Preserve the event, partition, offset, failure category, and schema information needed for investigation.

    Retry Topics and Dead-letter Topics

    Some consumer designs publish failed records to retry topics or dead-letter topics rather than retrying indefinitely in the primary partition loop.

    Primary Topic
          |
          v
    Consumer processing
          |
          +-- Success:
          |      commit progress
          |
          +-- Retryable failure:
          |      publish to retry path
          |
          +-- Permanent failure:
                 publish to dead-letter
                 handling path

    Moving an event out of the original partition can change ordering. Apply this pattern only when the business workflow defines how later events may proceed.

    Backpressure

    Kafka retains records while consumers progress at their own speed, but retention is not unlimited consumer capacity.

    Consumer-side controls include:

    • Bounded batch size
    • Bounded processing concurrency
    • Partition pause and resume
    • Consumer autoscaling within partition limits
    • Downstream connection budgets
    • Producer rate limiting
    • Optional-event load shedding

    Downstream Bottlenecks

    Increasing consumer count can overload the database or service used by the consumer.

    Kafka consumers:
    
    20 instances
    
    
    Database connection limit:
    
    50 connections
    
    
    Each consumer opens:
    
    5 connections
    
    
    Potential demand:
    
    100 connections

    Consumer scaling must respect downstream connection, CPU, transaction, storage, and API limits.

    Pause and Resume

    A consumer can temporarily pause selected assigned partitions when local or downstream processing capacity is unavailable, where supported by the consumer client.

    Downstream service slows
          |
          v
    Consumer pauses partition fetch
          |
          v
    Complete bounded in-flight work
          |
          v
    Downstream recovers
          |
          v
    Consumer resumes partition fetch

    Pausing the consumer does not stop producers from appending events. Consumer lag continues growing while the partition remains paused.

    Conceptual Consumer Configuration

    consumer:
      bootstrapServers:
        - approved-kafka-bootstrap-address
    
      groupId: learning-progress-projection
    
      subscription:
        topics:
          - learning-domain-events
    
      deserialization:
        key: approved-key-deserializer
        value: approved-value-deserializer
    
      offsetManagement:
        automaticCommit: false
        commitAfterDurableProcessing: true
    
      assignment:
        strategy: approved-assignment-strategy
    
      processing:
        maximumPollRecords: approved-batch-size
        boundedConcurrency: true
        idempotencyRequired: true
    
      retries:
        boundedAttempts: true
        exponentialBackoff: true
        deadLetterPath: approved-failure-topic
    
      observability:
        consumerLag: enabled
        oldestEventAge: enabled
        rebalanceEvents: enabled
        processingFailures: enabled
        offsetCommitFailures: enabled

    This configuration is conceptual. Exact property names, defaults, and guarantees depend on the Kafka version, consumer client, group protocol, and deployment platform.

    Java Consumer-group Example

    Properties properties = new Properties();
    
    properties.put(
        "bootstrap.servers",
        "kafka-bootstrap-address"
    );
    
    properties.put(
        "group.id",
        "learning-progress-projection"
    );
    
    properties.put(
        "key.deserializer",
        "org.apache.kafka.common.serialization.StringDeserializer"
    );
    
    properties.put(
        "value.deserializer",
        "org.apache.kafka.common.serialization.StringDeserializer"
    );
    
    properties.put(
        "enable.auto.commit",
        "false"
    );
    
    KafkaConsumer<String, String> consumer =
        new KafkaConsumer<>(properties);
    
    consumer.subscribe(
        List.of(
            "learning-domain-events"
        )
    );
    
    try {
        while (true) {
            ConsumerRecords<String, String> records =
                consumer.poll(
                    Duration.ofSeconds(1)
                );
    
            for (
                ConsumerRecord<String, String> record
                : records
            ) {
                processIdempotently(
                    record.key(),
                    record.value(),
                    record.topic(),
                    record.partition(),
                    record.offset()
                );
            }
    
            consumer.commitSync();
        }
    } finally {
        consumer.close();
    }

    This is a simplified educational example. Production code requires schema validation, exception classification, partition-aware recovery, secure configuration, transaction handling, controlled retries, shutdown handling, and protection against committing offsets for unsuccessfully processed records.

    Offset-aware Processing Record

    {
      "consumerGroup": "learning-progress-projection",
      "topic": "learning-domain-events",
      "partition": 2,
      "offset": 8500,
      "eventId": "stable-unique-event-id",
      "processingResult": "success"
    }

    Topic, partition, and offset provide useful processing context. A stable business event ID remains valuable for deduplication across republishing or migration scenarios.

    Inspecting a Consumer Group

    bin/kafka-consumer-groups.sh \
      --bootstrap-server kafka-bootstrap-address \
      --describe \
      --group learning-progress-projection

    Consumer-group inspection commonly includes topic, partition, current offset, log-end offset, lag, and active consumer-assignment information.

    Queue-like and Publish-subscribe Behaviour

    Consumer groups allow Kafka to provide both work sharing and independent subscriptions.

    Queue-like Work Sharing

    Same group ID:
    
    Consumer A
    Consumer B
    Consumer C
    
    
    Result:
    
    Consumers divide partitions
    and share processing.

    Publish-subscribe Behaviour

    Different group IDs:
    
    Notification Group
    Analytics Group
    Search Group
    
    
    Result:
    
    Each group independently
    reads the topic.

    Scaling a Consumer Group

    Scale a consumer group after identifying the actual throughput constraint.

    Condition Possible Direction
    Consumer CPU is saturated and unassigned partitions exist Add consumer instances within the partition limit.
    Every partition already has one consumer Optimize processing or increase partitions through a controlled design.
    One partition is hot Improve key distribution or divide safely independent work.
    Database is saturated Do not add consumers blindly; control downstream concurrency.
    Consumers spend time waiting on remote services Review batching, asynchronous I/O, timeouts, and downstream capacity.
    Rebalances are frequent Stabilize membership, processing time, and deployment behaviour.

    Increasing Partition Count

    Increasing topic partitions can increase future consumer-group parallelism. It also changes partition placement for newly produced records under many default key-to-partition calculations.

    Before adding partitions, evaluate:

    • Existing key-ordering requirements
    • Producer partitioning behaviour
    • Consumer scaling requirements
    • Broker storage and replication capacity
    • Hot-key distribution
    • Operational and monitoring overhead

    Partition rule: Partition count is a capacity and ordering decision, not merely a consumer-count setting. Changing it can affect how future keyed records are distributed.

    Learning-platform Example

    Kafka Topic:
    
    learning-domain-events
    
    
    Partitions selected by:
    
    learnerId
    
    
    Consumer Group 1:
    
    progress-projection
    
    
    Consumer Group 2:
    
    notification-service
    
    
    Consumer Group 3:
    
    learning-analytics

    All three groups can read the same retained events independently.

    Progress-projection Group

    Partition 0 -> Progress Consumer A
    
    Partition 1 -> Progress Consumer B
    
    Partition 2 -> Progress Consumer C
    
    Partition 3 -> Progress Consumer D

    Events for one learner remain in one partition when the producer uses a stable learner key. Different learners can be processed in parallel across partitions.

    Example Consumer Groups

    Consumer Group Purpose Processing Requirement
    Progress projection Build the current learner-progress view Idempotent updates with per-learner ordering
    Notification service Create notification commands from relevant events Prevent duplicate notification tasks
    Search indexing Update the searchable course projection Support replay when the index is rebuilt
    Learning analytics Process events for analytical models Track lag and tolerate a documented processing delay
    Certificate eligibility Evaluate qualifying completion events Generate one certificate outcome per valid completion

    Security Considerations

    Consumer groups should have only the permissions required by their applications.

    • Restrict topic read access by application responsibility
    • Restrict group access to approved group identifiers
    • Protect authentication credentials and certificates
    • Encrypt connections according to organizational policy
    • Use trusted tenant context from event data and authorization policy
    • Do not expose sensitive payload fields in consumer logs
    • Audit topic, group, and access-control changes
    • Separate sensitive event streams where required

    Observability

    Useful consumer-group metrics include:

    • Current offset by topic and partition
    • Log-end offset by partition
    • Consumer lag by partition
    • Oldest unprocessed-event age
    • Records processed per second
    • Processing duration by event type
    • Offset-commit success and failure
    • Consumer-group member count
    • Assigned partitions per consumer
    • Rebalance count and duration
    • Processing failure count
    • Retry and dead-letter volume
    • Consumer CPU, memory, and network usage
    • Downstream database or API latency
    • Hot-partition traffic

    Structured Consumer Event

    {
      "consumerGroup": "progress-projection",
      "topic": "learning-domain-events",
      "partition": 2,
      "offset": 8500,
      "processingResult": "success",
      "attempt": 1,
      "schemaVersion": 1
    }

    Avoid logging complete private payloads, authentication credentials, access tokens, or reusable secrets.

    Alert Conditions

    Alert when:

    • Consumer lag continues growing
    • The oldest unprocessed-event age exceeds the processing objective
    • One partition has substantially more lag than peer partitions
    • No active consumer owns a required partition
    • Offset commits repeatedly fail
    • Consumer-group rebalances occur repeatedly
    • Consumer membership changes unexpectedly
    • Processing failures increase
    • Retry or dead-letter volume grows
    • A poison event prevents partition progress
    • Consumer completion rate remains below producer rate
    • A downstream database or service approaches saturation

    Troubleshooting Workflow

    1. Identify the consumer group and subscribed topics.
    2. Check the active member count.
    3. Check current partition assignments.
    4. Compare partition count with active consumers.
    5. Inspect committed and log-end offsets.
    6. Identify partitions with growing lag.
    7. Check event-production and consumer-processing rates.
    8. Check recent group rebalances.
    9. Check consumer polling and processing duration.
    10. Check offset-commit failures.
    11. Check poison events and retry loops.
    12. Check CPU, memory, network, and downstream dependencies.
    13. Check whether one partition key is creating a hotspot.
    14. Scale, optimize, pause, replay, or repair through the approved procedure.

    Common Consumer-group Mistakes

    1

    Giving Independent Applications the Same Group ID

    The applications divide partitions instead of each receiving the complete topic.

    2

    Giving Every Instance a Different Group ID

    Every instance independently processes the complete event stream, potentially duplicating intended work.

    3

    Adding More Consumers than Partitions

    Additional consumers remain without active partition assignments.

    4

    Committing Offsets before Processing

    A crash can skip a required business operation.

    5

    Ignoring Duplicate Processing

    A crash after business commit but before offset commit can repeat the event.

    6

    Assuming Global Ordering

    Kafka ordering applies within a partition, not across the complete topic.

    7

    Using a Hot Partition Key

    One consumer becomes overloaded while other consumers remain underused.

    8

    Processing Too Many Records per Poll

    The consumer can exceed its processing boundary before returning to normal polling.

    9

    Retrying a Poison Event Indefinitely

    One bad event can prevent later records in the partition from progressing.

    10

    Scaling Consumers without Checking the Database

    Additional consumers overload the downstream connection pool or database.

    11

    Monitoring Only Total Group Lag

    One hot partition can be hidden by low lag on other partitions.

    12

    Changing Partition Count without Evaluating Key Distribution

    Future keyed records can be assigned differently, affecting ordering and locality assumptions.

    Recommended Test Cases

    Test Expected Evidence
    One consumer The consumer receives all subscribed partition assignments.
    Add a consumer Partitions are redistributed among active members.
    Remove a consumer Remaining members receive the released partitions.
    More consumers than partitions Excess consumers remain without active partition assignments.
    Two different group IDs Both groups independently process the topic records.
    Consumer failure before commit A replacement consumer reprocesses from the last committed position.
    Failure after business commit Idempotency prevents a duplicate business outcome.
    Hot partition Partition-level lag exposes the uneven workload.
    Poison event Bounded retries route the event through the approved failure path.
    Long processing The consumer remains stable or the work is moved to a more suitable asynchronous path.
    Consumer-group rebalance Revoked partitions stop processing before reassignment.
    Offset replay The projection can be rebuilt without uncontrolled duplicate side effects.
    Downstream slowdown Bounded concurrency protects the database or service.
    Schema evolution Old and new consumers remain compatible during deployment.

    Consumer-group Best Practices

    Recommended Practices

    • Use one stable group ID for instances of the same consumer application.
    • Use different group IDs for independent applications.
    • Choose partition count from throughput, ordering, and scaling requirements.
    • Use a stable business key for required per-entity ordering.
    • Avoid hot partition keys.
    • Do not add consumers beyond useful partition parallelism.
    • Commit offsets only after the intended durable processing boundary.
    • Make every consumer idempotent.
    • Keep batches within safe processing limits.
    • Use bounded retries with backoff.
    • Provide a controlled path for poison events.
    • Monitor lag by topic and partition.
    • Monitor oldest unprocessed-event age.
    • Monitor rebalances and membership changes.
    • Protect downstream databases with bounded concurrency.
    • Use pause and resume when temporary backpressure is required.
    • Evaluate static membership where stable instance identity is useful.
    • Test consumer joins, failures, restarts, and deployment rollouts.
    • Test offset replay and projection rebuilding.
    • Secure topic and group access using least privilege.

    Practice Exercise

    Design consumer groups for the event-processing workflows in your online learning platform.

    Requirements

    1. Create a topic with several partitions.
    2. Use learner ID as the partition key.
    3. Create a progress-projection consumer group.
    4. Create a separate analytics consumer group.
    5. Create a separate notification consumer group.
    6. Run one progress consumer and inspect its assignments.
    7. Add another consumer and observe partition redistribution.
    8. Run more consumers than partitions.
    9. Disable automatic offset commit.
    10. Commit offsets after durable processing.
    11. Add stable event identifiers and consumer deduplication.
    12. Stop one active consumer unexpectedly.
    13. Verify partition reassignment and redelivery behaviour.
    14. Introduce one poison event.
    15. Apply bounded retry and failure handling.
    16. Measure consumer lag by partition.
    17. Replay events to rebuild the progress projection.
    18. Test a downstream database slowdown.

    Consumer-group Design Template

    Decision Selected Direction Primary Risk Controlled
    Group ID Stable identifier per consumer application Prevents accidental work sharing between independent applications
    Partition key Stable learner or aggregate identifier Preserves required per-entity event ordering
    Partition count Capacity-tested partition count Provides useful consumer parallelism
    Offset commit After durable business processing Prevents silently skipped work
    Duplicate handling Transactional idempotency record Prevents repeated business effects
    Poison event Bounded retry and controlled failure topic Prevents one event from blocking partition progress indefinitely
    Scaling Consumers limited by partitions and downstream capacity Prevents idle consumers and dependency overload
    Monitoring Partition lag, event age, failures, and rebalances Detects delay and instability early

    Frequently Asked Questions

    1

    What is a Kafka consumer group?

    A consumer group is a set of consumers that cooperate to process records from one or more topics by dividing partition assignments.

    2

    What is a group ID?

    A group ID is the identifier that tells Kafka which consumers belong to the same logical consumer application.

    3

    Can two consumers in one group read the same partition?

    One partition is assigned to one active consumer within a consumer group at a given time.

    4

    Can different groups read the same event?

    Yes. Different groups maintain independent offsets and can process the same retained topic records.

    5

    What happens when consumers outnumber partitions?

    Some consumers remain without active partition assignments.

    6

    What is an offset?

    An offset identifies a record position inside a Kafka partition.

    7

    What is an offset commit?

    An offset commit records the consumer group's processing progress for a partition.

    8

    What is a consumer-group rebalance?

    A rebalance redistributes topic partitions after group membership, subscription, or relevant topic metadata changes.

    9

    What is consumer lag?

    Consumer lag is the difference between the latest partition position and the consumer group's current committed position.

    10

    Why can an event be processed twice?

    The business update can succeed while the offset commit fails or is not completed before a consumer failure.

    11

    Does adding consumers always increase throughput?

    No. Throughput is limited by available partitions, hot partitions, consumer resources, event cost, and downstream-system capacity.

    12

    How should a consumer group be scaled?

    Scale by measuring partition-level lag and processing capacity, then add consumers only when useful partitions and downstream capacity are available.

    Key Takeaway

    Kafka consumer groups allow instances of one application to process topic partitions cooperatively. Consumers sharing a group ID divide partition assignments, while applications using different group IDs independently consume the same retained event stream. Partition count determines the group's maximum useful partition-processing parallelism, and ordering is guaranteed within each partition rather than across the full topic. Consumer offsets track processing progress and support restart, recovery, and replay. Commit offsets only after the intended durable business operation, and make consumers idempotent because failure after processing but before commit can cause redelivery. Monitor lag and event age by partition, control poison events, stabilize group membership, protect downstream databases, and do not assume that adding consumers can divide one hot partition. A reliable consumer-group design aligns group IDs, partition keys, partition counts, offset management, idempotency, backpressure, retries, security, and observability with the business workflow.