consumer groups
Consumer Groups
Learn how Kafka consumer groups allow several consumers to process topic partitions in parallel, how group IDs create independent subscriptions, how partitions are assigned and rebalanced, how offsets track progress, and how lag, duplicate processing, failures, scaling, hot partitions, poison events, retries, static membership, idempotency, and observability affect production consumers.
Introduction
A Kafka topic can contain a continuous stream of events. One consumer can process that stream, but a single consumer eventually reaches limits in CPU, memory, network, database access, or event-processing capacity.
A consumer group allows several consumer instances belonging to the same application to cooperate.
Kafka Topic
|
+-- Partition 0
+-- Partition 1
+-- Partition 2
+-- Partition 3
|
v
Consumer Group
|
+-- Consumer A
+-- Consumer B
Kafka assigns topic partitions among the active members of the group. Each partition is processed by one active consumer within that group at a time.
Core idea: Consumers in the same group divide partition processing. Consumers in different groups independently read the same retained topic records.
Prerequisites
| # | Prerequisite | Why It Is Needed |
|---|---|---|
| 1 | Kafka topics and partitions | Partitions are the units distributed among consumers in a group. |
| 2 | Kafka producers and consumers | Producers append records, while consumers retrieve and process them. |
| 3 | Offsets | A consumer group tracks its progress separately for every partition. |
| 4 | Idempotency | Failures and offset recovery can cause an event to be processed again. |
| 5 | Partition ordering | Kafka ordering is scoped to an individual partition. |
| 6 | Backpressure | Consumer groups can fall behind when event production exceeds processing capacity. |
| 7 | Observability | Consumer lag, processing failures, offset commits, and rebalances require monitoring. |
What Is a Consumer Group?
A Kafka consumer group is a collection of consumers that work together to process records from one or more topics.
Every consumer in the group uses the same group identifier, commonly
configured through group.id.
group.id = certificate-processors
Members:
certificate-worker-1
certificate-worker-2
certificate-worker-3
Kafka assigns partitions among the group members so that the members can process different partitions concurrently without intentionally repeating the same partition work inside the group.
Group ID
The group ID identifies the logical consumer application.
Topic:
learning-domain-events
Consumer Group:
notification-service
Consumers:
notification-1
notification-2
notification-3
Consumers using the same group ID cooperate and divide available partitions.
Consumers using different group IDs represent independent subscriptions.
Topic:
learning-domain-events
Group 1:
notification-service
Group 2:
analytics-service
Group 3:
search-index-service
Each group can process the complete topic independently and maintain its own offsets.
Group-ID rule: Use the same group ID when application instances should share processing. Use different group IDs when applications need independent copies of the event stream.
Same Group vs Different Groups
| Configuration | Behaviour |
|---|---|
| Same group ID | Consumers cooperate and divide partition ownership. |
| Different group IDs | Consumers read the topic independently and maintain separate offsets. |
| One consumer, one group | The consumer processes all assigned topic partitions. |
| Several consumers, one group | Partitions are distributed among active group members. |
| Several consumer groups | Every group can independently process the retained stream. |
Partitions Determine Parallelism
Kafka distributes partitions, not individual records, among active members of a consumer group.
Four Partitions and One Consumer
Partition 0 --+
Partition 1 --+
Partition 2 --+--> Consumer A
Partition 3 --+
One consumer processes all four partitions.
Four Partitions and Two Consumers
Partition 0 --+
+--> Consumer A
Partition 1 --+
Partition 2 --+
+--> Consumer B
Partition 3 --+
Four Partitions and Four Consumers
Partition 0 -> Consumer A
Partition 1 -> Consumer B
Partition 2 -> Consumer C
Partition 3 -> Consumer D
Four Partitions and Six Consumers
Partition 0 -> Consumer A
Partition 1 -> Consumer B
Partition 2 -> Consumer C
Partition 3 -> Consumer D
Consumer E -> No partition
Consumer F -> No partition
Consumers without partition assignments remain inactive for ordinary record processing in that group.
Maximum Active Consumers
For a group consuming one topic, a simplified upper bound for active partition-processing consumers is:
\[ MaximumActiveConsumers = NumberOfTopicPartitions \]
When a group subscribes to several topics, applicable partition assignments depend on the combined subscribed partitions and the selected assignment strategy.
Adding consumers beyond available partition assignments improves neither active parallelism nor throughput.
Ordering within a Consumer Group
Kafka preserves the order of records within one partition.
Partition for learner-1042:
Offset 100 -> AccountRegistered
Offset 101 -> CourseEnrolled
Offset 102 -> LessonCompleted
Offset 103 -> CourseCompleted
The consumer assigned to this partition reads the records in partition order.
Records in separate partitions do not have one automatic global order.
Partition 0:
Event A1
Event A2
Event A3
Partition 1:
Event B1
Event B2
Event B3
No automatic total order exists
between Event A2 and Event B2.
Partition Key and Group Processing
A producer commonly uses a business key to place related records in the same partition.
Kafka record key:
learner-1042
Related events:
LearnerRegistered
LearnerEnrolled
LessonCompleted
CourseCompleted
Result:
Events can be routed to
one partition and processed
in partition order.
A poor key can create a hot partition, while a random key can break required per-entity ordering.
Consumer Offsets
An offset identifies a record's position inside a partition.
Partition 2:
Offset 8,000 -> Event A
Offset 8,001 -> Event B
Offset 8,002 -> Event C
Each consumer group maintains its own offset progress for each partition.
Topic Partition 2:
Notification Group:
Committed offset = 8,002
Analytics Group:
Committed offset = 7,950
Search Group:
Committed offset = 7,700
One slow consumer group does not change the positions of the other groups.
Offset Commit
An offset commit records the consumer group's processing position.
Consumer fetches:
Offsets 100 to 109
Consumer processes:
Offsets 100 to 109
Consumer commits:
Next required offset = 110
After restarting or receiving the partition again, a consumer can resume from the group's last committed position.
Auto Commit vs Manual Commit
| Mode | General Behaviour | Main Risk |
|---|---|---|
| Automatic offset commit | The client commits offsets according to configured automatic behaviour. | Progress can be committed before the complete business effect is durable if processing is not aligned correctly. |
| Manual synchronous commit | The consumer requests an offset commit and waits for its result. | Increases commit-path latency. |
| Manual asynchronous commit | The consumer requests a commit without blocking for the complete response path. | Commit failures and callback ordering require careful handling. |
Commit rule: Commit progress only after the application reaches the intended durable processing boundary.
Commit before Processing
Consumer fetches event
|
v
Consumer commits offset
|
v
Consumer begins processing
|
v
Consumer crashes
|
v
Restart begins after event
|
v
Business operation is skipped.
Committing too early can produce an at-most-once style failure in which required work is not completed.
Commit after Processing
Consumer fetches event
|
v
Consumer processes event
|
v
Business transaction commits
|
v
Consumer crashes before
offset commit
|
v
Event is delivered again.
Committing after processing protects against silently skipped work but can produce duplicate processing. The consumer must be idempotent.
Idempotent Consumer
An idempotent consumer produces one intended business outcome even when the same event is delivered more than once.
{
"eventId": "stable-unique-event-id",
"eventType": "LearnerEnrolled",
"eventVersion": 1,
"tenantId": 17,
"learnerId": 1042,
"courseId": 42
}
The event identifier can be recorded in the same transaction as the consumer's business update.
Deduplication Table
CREATE TABLE processed_consumer_events
(
consumer_group VARCHAR(150) NOT NULL,
event_id VARCHAR(150) NOT NULL,
processed_at TIMESTAMP NOT NULL,
PRIMARY KEY
(
consumer_group,
event_id
)
);
The unique constraint prevents the same consumer group from applying the same event twice when the check and business operation use an appropriate transactional design.
Rebalancing
Rebalancing redistributes partition assignments among active members of a consumer group.
A rebalance can occur when:
- A consumer joins the group
- A consumer leaves the group
- A consumer fails or becomes unresponsive
- The subscribed topic's partition count changes
- The group's topic subscriptions change
- Group membership or assignment configuration changes
Before:
Consumer A -> Partitions 0 and 1
Consumer B -> Partitions 2 and 3
Consumer C joins
|
v
Group rebalances
|
v
Possible result:
Consumer A -> Partitions 0 and 1
Consumer B -> Partition 2
Consumer C -> Partition 3
Consumer Failure
Consumer B owns:
Partitions 2 and 3
Consumer B stops responding
|
v
Group membership detects failure
|
v
Partitions are reassigned
|
v
Another consumer resumes
from committed offsets.
Records processed after the last committed offset can be processed again by the replacement consumer.
Heartbeats and Session Activity
Consumers use group-management communication to keep their membership active. Heartbeats help the group determine whether members are still participating.
Consumer
|
+-- Fetch records
+-- Process records
+-- Maintain group activity
+-- Commit offsets
If the group concludes that a member is unavailable, the member's partitions can be reassigned.
Slow Processing and Polling
Long processing operations can interfere with healthy group participation when the consumer does not interact with Kafka within its configured processing boundaries.
Consumer polls records
|
v
One record takes a long time
|
v
Consumer does not poll again
within expected boundary
|
v
Group can consider the member
unable to continue normally
|
v
Rebalance can occur.
Possible design directions include:
- Reduce records returned per poll
- Bound event-processing duration
- Move long-running work to a separate task queue
- Pause assigned partitions while bounded work completes
- Adjust consumer configuration using measured processing behaviour
Rebalance Impact
During partition reassignment, consumers must stop processing revoked partitions and begin processing newly assigned partitions from valid positions.
Frequent rebalances can cause:
- Temporary processing interruptions
- Repeated initialization work
- Cache rebuilding inside consumers
- Duplicate processing around offset boundaries
- Increased consumer lag
- Reduced overall throughput
Partition-assignment Strategies
A partition-assignment strategy determines how partitions are distributed among group members.
| Strategy Direction | General Goal |
|---|---|
| Range-style assignment | Assign contiguous topic-partition ranges to consumers. |
| Round-robin assignment | Distribute subscribed partitions across consumers in rotation. |
| Sticky assignment | Balance partitions while trying to preserve prior assignments. |
| Cooperative sticky assignment | Move affected assignments incrementally rather than revoking all ownership at once where supported. |
Final selection depends on subscription patterns, Kafka client support, deployment behaviour, balance requirements, and the cost of partition movement.
Static Membership
Dynamic consumers can receive new generated member identities when they restart. Static membership allows stable consumer identities to be associated with application instances where supported and configured.
Stable membership can help reduce unnecessary partition movement during controlled short restarts, but failed or duplicated member identities still require correct lifecycle handling.
Consumer instance:
group.id = search-indexers
group.instance.id = search-indexer-03
Consumer Lag
Consumer lag measures how far a consumer group remains behind the latest available partition position.
A simplified record-based calculation is:
\[ ConsumerLag = LogEndOffset - CurrentCommittedOffset \]
Partition log-end offset:
50,000
Consumer-group offset:
47,500
Consumer lag:
2,500 records
Lag should be measured per partition. One heavily loaded partition can fall behind even when the group's total or average lag appears acceptable.
Event-age Lag
Record count does not directly express business delay because events can differ in size and processing cost.
Group A:
10,000 small events behind
Group B:
100 expensive events behind
Monitor the age of the oldest unprocessed event as well as offset distance.
Lag Growth
Let \(P\) represent the event-production rate and \(C\) represent the consumer group's successful processing rate.
Lag grows when:
\[ P > C \]
A simplified lag-growth rate is:
\[ LagGrowthRate = P - C \]
A group can drain its backlog only when its sustained completion rate exceeds the incoming production rate for the affected partitions.
Estimated Drain Time
When consumer capacity exceeds incoming production, a simplified estimate is:
\[ EstimatedDrainTime = \frac{ CurrentLag }{ ConsumerRate - ProducerRate } \]
This is a planning approximation. Event cost, retries, partition skew, consumer rebalances, and downstream dependencies affect actual recovery.
Hot Partitions
Adding group members cannot divide one Kafka partition among several active consumers in the same group.
Topic has four partitions:
Partition 0:
5,000 events per second
Partition 1:
100 events per second
Partition 2:
120 events per second
Partition 3:
90 events per second
The consumer assigned to Partition 0 becomes the group's bottleneck. Consumers assigned to other partitions can remain underused.
Possible controls include:
- Choose a more evenly distributed producer key
- Split safely independent work across controlled subkeys
- Increase partitions before current capacity is exhausted
- Separate the heavy workload into another topic
- Optimize processing for the hot event class
- Apply producer-side rate or admission limits
Scaling rule: Adding consumers helps only when Kafka can assign additional partitions containing useful work to those consumers. It does not divide one hot partition automatically.
Poison Events
A poison event repeatedly fails because of invalid data, unsupported schema, a code defect, or an unavailable dependency.
Consumer reads Offset 500
|
v
Processing fails
|
v
Consumer retries Offset 500
|
v
Processing fails again
|
v
Later partition records
cannot progress safely.
Use bounded retries and a controlled failure path. Preserve the event, partition, offset, failure category, and schema information needed for investigation.
Retry Topics and Dead-letter Topics
Some consumer designs publish failed records to retry topics or dead-letter topics rather than retrying indefinitely in the primary partition loop.
Primary Topic
|
v
Consumer processing
|
+-- Success:
| commit progress
|
+-- Retryable failure:
| publish to retry path
|
+-- Permanent failure:
publish to dead-letter
handling path
Moving an event out of the original partition can change ordering. Apply this pattern only when the business workflow defines how later events may proceed.
Backpressure
Kafka retains records while consumers progress at their own speed, but retention is not unlimited consumer capacity.
Consumer-side controls include:
- Bounded batch size
- Bounded processing concurrency
- Partition pause and resume
- Consumer autoscaling within partition limits
- Downstream connection budgets
- Producer rate limiting
- Optional-event load shedding
Downstream Bottlenecks
Increasing consumer count can overload the database or service used by the consumer.
Kafka consumers:
20 instances
Database connection limit:
50 connections
Each consumer opens:
5 connections
Potential demand:
100 connections
Consumer scaling must respect downstream connection, CPU, transaction, storage, and API limits.
Pause and Resume
A consumer can temporarily pause selected assigned partitions when local or downstream processing capacity is unavailable, where supported by the consumer client.
Downstream service slows
|
v
Consumer pauses partition fetch
|
v
Complete bounded in-flight work
|
v
Downstream recovers
|
v
Consumer resumes partition fetch
Pausing the consumer does not stop producers from appending events. Consumer lag continues growing while the partition remains paused.
Conceptual Consumer Configuration
consumer:
bootstrapServers:
- approved-kafka-bootstrap-address
groupId: learning-progress-projection
subscription:
topics:
- learning-domain-events
deserialization:
key: approved-key-deserializer
value: approved-value-deserializer
offsetManagement:
automaticCommit: false
commitAfterDurableProcessing: true
assignment:
strategy: approved-assignment-strategy
processing:
maximumPollRecords: approved-batch-size
boundedConcurrency: true
idempotencyRequired: true
retries:
boundedAttempts: true
exponentialBackoff: true
deadLetterPath: approved-failure-topic
observability:
consumerLag: enabled
oldestEventAge: enabled
rebalanceEvents: enabled
processingFailures: enabled
offsetCommitFailures: enabled
This configuration is conceptual. Exact property names, defaults, and guarantees depend on the Kafka version, consumer client, group protocol, and deployment platform.
Java Consumer-group Example
Properties properties = new Properties();
properties.put(
"bootstrap.servers",
"kafka-bootstrap-address"
);
properties.put(
"group.id",
"learning-progress-projection"
);
properties.put(
"key.deserializer",
"org.apache.kafka.common.serialization.StringDeserializer"
);
properties.put(
"value.deserializer",
"org.apache.kafka.common.serialization.StringDeserializer"
);
properties.put(
"enable.auto.commit",
"false"
);
KafkaConsumer<String, String> consumer =
new KafkaConsumer<>(properties);
consumer.subscribe(
List.of(
"learning-domain-events"
)
);
try {
while (true) {
ConsumerRecords<String, String> records =
consumer.poll(
Duration.ofSeconds(1)
);
for (
ConsumerRecord<String, String> record
: records
) {
processIdempotently(
record.key(),
record.value(),
record.topic(),
record.partition(),
record.offset()
);
}
consumer.commitSync();
}
} finally {
consumer.close();
}
This is a simplified educational example. Production code requires schema validation, exception classification, partition-aware recovery, secure configuration, transaction handling, controlled retries, shutdown handling, and protection against committing offsets for unsuccessfully processed records.
Offset-aware Processing Record
{
"consumerGroup": "learning-progress-projection",
"topic": "learning-domain-events",
"partition": 2,
"offset": 8500,
"eventId": "stable-unique-event-id",
"processingResult": "success"
}
Topic, partition, and offset provide useful processing context. A stable business event ID remains valuable for deduplication across republishing or migration scenarios.
Inspecting a Consumer Group
bin/kafka-consumer-groups.sh \
--bootstrap-server kafka-bootstrap-address \
--describe \
--group learning-progress-projection
Consumer-group inspection commonly includes topic, partition, current offset, log-end offset, lag, and active consumer-assignment information.
Queue-like and Publish-subscribe Behaviour
Consumer groups allow Kafka to provide both work sharing and independent subscriptions.
Queue-like Work Sharing
Same group ID:
Consumer A
Consumer B
Consumer C
Result:
Consumers divide partitions
and share processing.
Publish-subscribe Behaviour
Different group IDs:
Notification Group
Analytics Group
Search Group
Result:
Each group independently
reads the topic.
Scaling a Consumer Group
Scale a consumer group after identifying the actual throughput constraint.
| Condition | Possible Direction |
|---|---|
| Consumer CPU is saturated and unassigned partitions exist | Add consumer instances within the partition limit. |
| Every partition already has one consumer | Optimize processing or increase partitions through a controlled design. |
| One partition is hot | Improve key distribution or divide safely independent work. |
| Database is saturated | Do not add consumers blindly; control downstream concurrency. |
| Consumers spend time waiting on remote services | Review batching, asynchronous I/O, timeouts, and downstream capacity. |
| Rebalances are frequent | Stabilize membership, processing time, and deployment behaviour. |
Increasing Partition Count
Increasing topic partitions can increase future consumer-group parallelism. It also changes partition placement for newly produced records under many default key-to-partition calculations.
Before adding partitions, evaluate:
- Existing key-ordering requirements
- Producer partitioning behaviour
- Consumer scaling requirements
- Broker storage and replication capacity
- Hot-key distribution
- Operational and monitoring overhead
Partition rule: Partition count is a capacity and ordering decision, not merely a consumer-count setting. Changing it can affect how future keyed records are distributed.
Learning-platform Example
Kafka Topic:
learning-domain-events
Partitions selected by:
learnerId
Consumer Group 1:
progress-projection
Consumer Group 2:
notification-service
Consumer Group 3:
learning-analytics
All three groups can read the same retained events independently.
Progress-projection Group
Partition 0 -> Progress Consumer A
Partition 1 -> Progress Consumer B
Partition 2 -> Progress Consumer C
Partition 3 -> Progress Consumer D
Events for one learner remain in one partition when the producer uses a stable learner key. Different learners can be processed in parallel across partitions.
Example Consumer Groups
| Consumer Group | Purpose | Processing Requirement |
|---|---|---|
| Progress projection | Build the current learner-progress view | Idempotent updates with per-learner ordering |
| Notification service | Create notification commands from relevant events | Prevent duplicate notification tasks |
| Search indexing | Update the searchable course projection | Support replay when the index is rebuilt |
| Learning analytics | Process events for analytical models | Track lag and tolerate a documented processing delay |
| Certificate eligibility | Evaluate qualifying completion events | Generate one certificate outcome per valid completion |
Security Considerations
Consumer groups should have only the permissions required by their applications.
- Restrict topic read access by application responsibility
- Restrict group access to approved group identifiers
- Protect authentication credentials and certificates
- Encrypt connections according to organizational policy
- Use trusted tenant context from event data and authorization policy
- Do not expose sensitive payload fields in consumer logs
- Audit topic, group, and access-control changes
- Separate sensitive event streams where required
Observability
Useful consumer-group metrics include:
- Current offset by topic and partition
- Log-end offset by partition
- Consumer lag by partition
- Oldest unprocessed-event age
- Records processed per second
- Processing duration by event type
- Offset-commit success and failure
- Consumer-group member count
- Assigned partitions per consumer
- Rebalance count and duration
- Processing failure count
- Retry and dead-letter volume
- Consumer CPU, memory, and network usage
- Downstream database or API latency
- Hot-partition traffic
Structured Consumer Event
{
"consumerGroup": "progress-projection",
"topic": "learning-domain-events",
"partition": 2,
"offset": 8500,
"processingResult": "success",
"attempt": 1,
"schemaVersion": 1
}
Avoid logging complete private payloads, authentication credentials, access tokens, or reusable secrets.
Alert Conditions
Alert when:
- Consumer lag continues growing
- The oldest unprocessed-event age exceeds the processing objective
- One partition has substantially more lag than peer partitions
- No active consumer owns a required partition
- Offset commits repeatedly fail
- Consumer-group rebalances occur repeatedly
- Consumer membership changes unexpectedly
- Processing failures increase
- Retry or dead-letter volume grows
- A poison event prevents partition progress
- Consumer completion rate remains below producer rate
- A downstream database or service approaches saturation
Troubleshooting Workflow
- Identify the consumer group and subscribed topics.
- Check the active member count.
- Check current partition assignments.
- Compare partition count with active consumers.
- Inspect committed and log-end offsets.
- Identify partitions with growing lag.
- Check event-production and consumer-processing rates.
- Check recent group rebalances.
- Check consumer polling and processing duration.
- Check offset-commit failures.
- Check poison events and retry loops.
- Check CPU, memory, network, and downstream dependencies.
- Check whether one partition key is creating a hotspot.
- Scale, optimize, pause, replay, or repair through the approved procedure.
Common Consumer-group Mistakes
Giving Independent Applications the Same Group ID
The applications divide partitions instead of each receiving the complete topic.
Giving Every Instance a Different Group ID
Every instance independently processes the complete event stream, potentially duplicating intended work.
Adding More Consumers than Partitions
Additional consumers remain without active partition assignments.
Committing Offsets before Processing
A crash can skip a required business operation.
Ignoring Duplicate Processing
A crash after business commit but before offset commit can repeat the event.
Assuming Global Ordering
Kafka ordering applies within a partition, not across the complete topic.
Using a Hot Partition Key
One consumer becomes overloaded while other consumers remain underused.
Processing Too Many Records per Poll
The consumer can exceed its processing boundary before returning to normal polling.
Retrying a Poison Event Indefinitely
One bad event can prevent later records in the partition from progressing.
Scaling Consumers without Checking the Database
Additional consumers overload the downstream connection pool or database.
Monitoring Only Total Group Lag
One hot partition can be hidden by low lag on other partitions.
Changing Partition Count without Evaluating Key Distribution
Future keyed records can be assigned differently, affecting ordering and locality assumptions.
Recommended Test Cases
| Test | Expected Evidence |
|---|---|
| One consumer | The consumer receives all subscribed partition assignments. |
| Add a consumer | Partitions are redistributed among active members. |
| Remove a consumer | Remaining members receive the released partitions. |
| More consumers than partitions | Excess consumers remain without active partition assignments. |
| Two different group IDs | Both groups independently process the topic records. |
| Consumer failure before commit | A replacement consumer reprocesses from the last committed position. |
| Failure after business commit | Idempotency prevents a duplicate business outcome. |
| Hot partition | Partition-level lag exposes the uneven workload. |
| Poison event | Bounded retries route the event through the approved failure path. |
| Long processing | The consumer remains stable or the work is moved to a more suitable asynchronous path. |
| Consumer-group rebalance | Revoked partitions stop processing before reassignment. |
| Offset replay | The projection can be rebuilt without uncontrolled duplicate side effects. |
| Downstream slowdown | Bounded concurrency protects the database or service. |
| Schema evolution | Old and new consumers remain compatible during deployment. |
Consumer-group Best Practices
Recommended Practices
- Use one stable group ID for instances of the same consumer application.
- Use different group IDs for independent applications.
- Choose partition count from throughput, ordering, and scaling requirements.
- Use a stable business key for required per-entity ordering.
- Avoid hot partition keys.
- Do not add consumers beyond useful partition parallelism.
- Commit offsets only after the intended durable processing boundary.
- Make every consumer idempotent.
- Keep batches within safe processing limits.
- Use bounded retries with backoff.
- Provide a controlled path for poison events.
- Monitor lag by topic and partition.
- Monitor oldest unprocessed-event age.
- Monitor rebalances and membership changes.
- Protect downstream databases with bounded concurrency.
- Use pause and resume when temporary backpressure is required.
- Evaluate static membership where stable instance identity is useful.
- Test consumer joins, failures, restarts, and deployment rollouts.
- Test offset replay and projection rebuilding.
- Secure topic and group access using least privilege.
Practice Exercise
Design consumer groups for the event-processing workflows in your online learning platform.
Requirements
- Create a topic with several partitions.
- Use learner ID as the partition key.
- Create a progress-projection consumer group.
- Create a separate analytics consumer group.
- Create a separate notification consumer group.
- Run one progress consumer and inspect its assignments.
- Add another consumer and observe partition redistribution.
- Run more consumers than partitions.
- Disable automatic offset commit.
- Commit offsets after durable processing.
- Add stable event identifiers and consumer deduplication.
- Stop one active consumer unexpectedly.
- Verify partition reassignment and redelivery behaviour.
- Introduce one poison event.
- Apply bounded retry and failure handling.
- Measure consumer lag by partition.
- Replay events to rebuild the progress projection.
- Test a downstream database slowdown.
Consumer-group Design Template
| Decision | Selected Direction | Primary Risk Controlled |
|---|---|---|
| Group ID | Stable identifier per consumer application | Prevents accidental work sharing between independent applications |
| Partition key | Stable learner or aggregate identifier | Preserves required per-entity event ordering |
| Partition count | Capacity-tested partition count | Provides useful consumer parallelism |
| Offset commit | After durable business processing | Prevents silently skipped work |
| Duplicate handling | Transactional idempotency record | Prevents repeated business effects |
| Poison event | Bounded retry and controlled failure topic | Prevents one event from blocking partition progress indefinitely |
| Scaling | Consumers limited by partitions and downstream capacity | Prevents idle consumers and dependency overload |
| Monitoring | Partition lag, event age, failures, and rebalances | Detects delay and instability early |
Frequently Asked Questions
What is a Kafka consumer group?
A consumer group is a set of consumers that cooperate to process records from one or more topics by dividing partition assignments.
What is a group ID?
A group ID is the identifier that tells Kafka which consumers belong to the same logical consumer application.
Can two consumers in one group read the same partition?
One partition is assigned to one active consumer within a consumer group at a given time.
Can different groups read the same event?
Yes. Different groups maintain independent offsets and can process the same retained topic records.
What happens when consumers outnumber partitions?
Some consumers remain without active partition assignments.
What is an offset?
An offset identifies a record position inside a Kafka partition.
What is an offset commit?
An offset commit records the consumer group's processing progress for a partition.
What is a consumer-group rebalance?
A rebalance redistributes topic partitions after group membership, subscription, or relevant topic metadata changes.
What is consumer lag?
Consumer lag is the difference between the latest partition position and the consumer group's current committed position.
Why can an event be processed twice?
The business update can succeed while the offset commit fails or is not completed before a consumer failure.
Does adding consumers always increase throughput?
No. Throughput is limited by available partitions, hot partitions, consumer resources, event cost, and downstream-system capacity.
How should a consumer group be scaled?
Scale by measuring partition-level lag and processing capacity, then add consumers only when useful partitions and downstream capacity are available.
Key Takeaway
Kafka consumer groups allow instances of one application to process topic partitions cooperatively. Consumers sharing a group ID divide partition assignments, while applications using different group IDs independently consume the same retained event stream. Partition count determines the group's maximum useful partition-processing parallelism, and ordering is guaranteed within each partition rather than across the full topic. Consumer offsets track processing progress and support restart, recovery, and replay. Commit offsets only after the intended durable business operation, and make consumers idempotent because failure after processing but before commit can cause redelivery. Monitor lag and event age by partition, control poison events, stabilize group membership, protect downstream databases, and do not assume that adding consumers can divide one hot partition. A reliable consumer-group design aligns group IDs, partition keys, partition counts, offset management, idempotency, backpressure, retries, security, and observability with the business workflow.