Queues vs logs
Queues vs Logs
Learn how message queues and durable event logs move data between distributed services, why queues model work to be completed while logs model retained facts, and how acknowledgments, offsets, retention, replay, consumer groups, partitions, ordering, delivery guarantees, backpressure, dead-letter handling, compaction, idempotency, and scaling influence the correct messaging design.
Introduction
Distributed services often need to communicate without waiting for every downstream operation to finish inside the original request.
Client submits an operation
|
v
Application validates request
|
v
Message is published
|
v
Application responds
|
v
Background consumer processes message
Asynchronous messaging separates the producer that creates work or events from the consumer that processes them.
Two common messaging models are:
- Message queues
- Append-only event logs
Both models carry messages between producers and consumers, but they answer an important question differently:
What should happen to a message
after one consumer processes it?
In the traditional queue model, consumers compete for work. One consumer processes a queued message, acknowledges it, and the broker can remove it according to its policy.
In the log model, reading does not remove the event. The event remains in the retained log, and each consumer group tracks its own reading position.
Core idea: A queue primarily represents work waiting to be completed. A log primarily represents an ordered history of events that independent consumers can read and replay.
Prerequisites
| # | Prerequisite | Why It Is Needed |
|---|---|---|
| 1 | Synchronous vs asynchronous communication | Messaging allows processing to continue outside the original request path. |
| 2 | Producer-consumer pattern | Producers publish messages and consumers process them. |
| 3 | Idempotency | A consumer can receive the same logical message more than once. |
| 4 | Retries and backoff | Temporary failures require controlled redelivery and retry handling. |
| 5 | Partitioning | Logs commonly use partitions as units of ordering and parallel consumption. |
| 6 | Backpressure and overload control | Producers can generate work faster than consumers can process it. |
| 7 | Observability | Queue backlog, message age, consumer lag, retries, and failures must be monitored. |
What Is a Message Queue?
A message queue stores messages that represent work waiting for an eligible consumer.
Producer
|
v
+-----------------------+
| Queue |
| Task 1 |
| Task 2 |
| Task 3 |
+-----------------------+
|
+-- Consumer A
+-- Consumer B
+-- Consumer C
Consumers connected to the same queue normally compete for messages. A particular message is assigned to one consumer at a time according to the broker's delivery policy.
Typical queue workloads include:
- Sending an email
- Generating a report
- Resizing an uploaded image
- Processing a video
- Executing an integration request
- Running a background database operation
- Dispatching a notification
What Is an Event Log?
An event log is an append-only sequence of records. New events are appended to the end of a partition, and retained events remain available according to a time-based, size-based, or compaction policy.
Partition 0:
Offset 0 -> CourseCreated
Offset 1 -> CoursePublished
Offset 2 -> LearnerEnrolled
Offset 3 -> ProgressUpdated
Offset 4 -> CertificateIssued
A consumer reads records in sequence and tracks its position, commonly called an offset.
Consumer Group A:
Current offset = 5
Consumer Group B:
Current offset = 2
Consumer Group C:
Current offset = 0
The consumer groups can read the same event history independently and at different speeds.
Core Difference
| Question | Queue | Log |
|---|---|---|
| What does a message represent? | Work that should be completed | A fact or event that occurred |
| Who receives it? | Usually one eligible consumer from a competing group | Every interested consumer group can read it independently |
| What happens after reading? | The message can be removed after successful acknowledgment | Reading does not remove the event from the retained log |
| How is progress tracked? | Delivery, lock, visibility, and acknowledgment state | Consumer position or offset |
| Can data be replayed? | Usually limited once successfully consumed, unless copied elsewhere | Yes, while the required offsets remain inside retention |
| Primary scaling model | Add competing workers | Add partitions and distribute partitions among consumers |
| Typical purpose | Task dispatch and work distribution | Event distribution, replay, integration, and stream processing |
Queue Processing Model
Queue contains:
Task A
Task B
Task C
Task D
Consumers:
Worker 1 receives Task A
Worker 2 receives Task B
Worker 3 receives Task C
Task D remains waiting.
Adding workers increases processing parallelism until another resource becomes the bottleneck.
Each task should normally be handled by one successful worker. If the worker fails before acknowledgment, the broker can make the task available for another delivery according to its policy.
Log Consumption Model
Event log:
Event A
Event B
Event C
Event D
Consumer Group: Notifications
Reads A, B, C, D
Consumer Group: Analytics
Reads A, B, C, D
Consumer Group: Search Index
Reads A, B, C, D
The groups read the same retained records without requiring the producer to write a separate copy for each group.
Competing Consumers
A queue commonly distributes different messages among several workers.
Queue
/ | \
/ | \
v v v
Worker A Worker B Worker C
One message is completed
by one successful worker.
This pattern is useful when workers provide equivalent processing capability and each task needs one successful execution.
Consumer Groups
A log consumer group combines queue-like competition with log-based retention.
Topic with four partitions:
Partition 0
Partition 1
Partition 2
Partition 3
Consumer Group A:
Consumer A1 -> Partition 0
Consumer A2 -> Partition 1
Consumer A3 -> Partition 2
Consumer A4 -> Partition 3
Consumers inside one group divide partition ownership. Another group can independently consume all four partitions for a different purpose.
Consumer-group rule: Consumers inside one group share the work for that group. Separate groups independently receive the retained event stream.
Queue Acknowledgments
An acknowledgment tells the queue broker that a delivered message was processed successfully.
Broker delivers message
|
v
Consumer processes message
|
+-- Success:
| acknowledge
| broker completes message
|
+-- Failure:
reject, retry,
or allow redelivery
Acknowledging before the business operation commits can lose work if the consumer fails after acknowledgment.
Acknowledging after the operation commits can produce redelivery if the acknowledgment is lost.
Consumers should therefore be idempotent.
Log Offsets
An offset identifies a record's position within a log partition.
Partition 2:
Offset 100 -> Event A
Offset 101 -> Event B
Offset 102 -> Event C
Consumer committed offset:
101
The consumer uses its committed position to determine where processing should continue after restart or reassignment.
Advancing the offset before processing completes can skip work after a failure. Advancing it after processing completes can cause the event to be processed again after an uncertain failure.
Replay
Replay means moving a log consumer's position backward and processing retained events again.
Current consumer offset:
10,000
Replay starting offset:
8,000
Consumer processes:
8,000
8,001
8,002
...
10,000
Replay can support:
- Rebuilding a search index
- Recreating a materialized view
- Reprocessing events after correcting a consumer defect
- Producing a new analytical projection
- Onboarding a new consumer from retained history
Replay is safe only when consumers can handle repeated events and external side effects are protected.
Retention
Logs retain records according to an explicit policy rather than deleting each event when one consumer reads it.
Retention can be determined by:
- Event age
- Total retained size
- Per-partition size
- Key-based compaction
- Compliance and business requirements
Replay is possible only while the required event range remains available.
Log Compaction
Compaction retains a selected latest record for each key rather than every historical record indefinitely.
Original log:
course:42 -> Draft
course:42 -> Reviewed
course:42 -> Published
Compacted representation:
course:42 -> Published
Compaction can support state reconstruction, but it is not equivalent to a complete audit history because older values can be removed according to the compaction policy.
Queue Deletion vs Log Retention
| Lifecycle Event | Queue | Log |
|---|---|---|
| Consumer reads message | Message becomes delivered or in flight | Event remains in the log |
| Consumer completes processing | Acknowledgment can complete or remove the message | Consumer advances its offset |
| Another independent consumer needs the same data | Usually needs another queue or routed copy | Uses another consumer group |
| Historical reprocessing | Requires retained copies or another archive | Reset offset if the records remain retained |
Ordering in Queues
A queue can provide first-in-first-out behaviour under a specific scope, but concurrency, retries, priorities, and redelivery can affect the order in which processing completes.
Queue delivery:
Task 1 -> Worker A
Task 2 -> Worker B
Processing duration:
Task 1 takes 10 seconds
Task 2 takes 1 second
Completion order:
Task 2
Task 1
Delivery order and completion order are different concepts.
Ordering in Logs
A log normally provides ordering within a partition.
Partition 0:
Offset 1
Offset 2
Offset 3
Offset 4
Events in different partitions do not automatically share one global order.
Partition A:
A1
A2
A3
Partition B:
B1
B2
B3
No automatic total order exists
between A2 and B2.
Events requiring ordered processing should use a partition key that places related events in the same partition.
Partition Keys
A log producer commonly supplies a key that determines the destination partition.
Event key:
learner-1042
All related progress events:
ProgressStarted
LessonCompleted
QuizCompleted
CourseCompleted
Route to the same partition.
This can preserve order for one learner while allowing events for other learners to be processed in parallel.
Hot Partitions
A partition key with highly uneven traffic can create a hot partition.
Normal course keys:
1,000 events per minute
One large course key:
500,000 events per minute
Result:
The partition owning that key
becomes overloaded.
Adding consumers cannot increase parallel processing for one partition beyond the broker and consumer model's partition-level limit.
Possible controls include:
- Selecting a higher-cardinality partition key
- Splitting safely independent workloads into subkeys
- Using controlled bucketing
- Adding partitions before capacity is exhausted
- Applying producer rate limits
- Separating heavy workloads into another topic
Queue Parallelism
Queue workers commonly draw messages from one waiting-work collection.
Queue backlog:
100,000 tasks
Workers:
10 workers
|
v
Workers compete for
available tasks.
Additional workers can increase throughput until the queue, database, storage, network, or downstream service reaches its safe concurrency limit.
Log Parallelism
Partition count commonly determines the maximum active partition-processing parallelism within one consumer group.
Topic:
4 partitions
Consumer group:
2 consumers
-> Each consumer handles
multiple partitions
Consumer group:
4 consumers
-> One active consumer
per partition
Consumer group:
6 consumers
-> Some consumers have
no active partition
Increasing consumers beyond the relevant partition count does not create additional active partition ownership in that group.
Consumer-group Rebalancing
When a log consumer joins, leaves, or fails, partition ownership can be reassigned among the remaining consumers.
Consumer Group:
Consumer A -> Partitions 0 and 1
Consumer B -> Partitions 2 and 3
Consumer B fails
|
v
Group rebalances
|
v
Consumer A or a replacement
receives Partitions 2 and 3
Consumers must stop using revoked partitions, resume from committed offsets, and handle repeated processing safely around reassignment.
Delivery Guarantees
| Guarantee | General Meaning | Application Concern |
|---|---|---|
| At-most-once | A message is processed zero or one time | Failures can cause work to be skipped |
| At-least-once | A message is retried until acknowledged or otherwise completed | Duplicate processing is possible |
| Effectively-once business outcome | Repeated delivery produces one intended business effect | Requires idempotency, deduplication, or transactional coordination |
Delivery rule: Do not assume that broker delivery automatically guarantees one business effect. Protect the consumer's state change against duplicate delivery and uncertain acknowledgments.
Idempotent Consumers
An idempotent consumer safely recognizes or absorbs repeated delivery of the same logical operation.
{
"eventId": "unique-event-id",
"eventType": "LearnerEnrolled",
"tenantId": 17,
"learnerId": 1042,
"courseId": 42
}
The consumer can record the event identifier in the same transactional boundary as the resulting business update.
Conceptual Deduplication Table
CREATE TABLE processed_messages
(
consumer_name VARCHAR(100) NOT NULL,
message_id VARCHAR(150) NOT NULL,
processed_at TIMESTAMP NOT NULL,
PRIMARY KEY
(
consumer_name,
message_id
)
);
Dead-letter Handling
A message that repeatedly fails should not block useful work indefinitely.
Message delivery
|
v
Processing fails
|
v
Bounded retry policy
|
+-- Later success:
| complete message
|
+-- Retry limit reached:
move to dead-letter
handling path
A dead-letter workflow should preserve:
- Message identity
- Failure reason
- Attempt count
- Original destination
- Relevant correlation context
- Safe replay or disposal procedure
Avoid including secrets or unnecessary private data in dead-letter metadata.
Poison Messages
A poison message consistently fails because of invalid data, unsupported format, missing dependency, or a consumer defect.
Poison message
|
v
Consumer fails
|
v
Immediate retry
|
v
Consumer fails again
|
v
Partition or queue progress
is repeatedly disrupted
Use bounded retries, backoff, dead-letter handling, and operator-visible diagnostics.
Backpressure
Backpressure communicates that consumers cannot safely keep pace with producers.
Producer rate:
10,000 messages per second
Consumer capacity:
4,000 messages per second
Difference:
6,000 messages per second
accumulate in backlog.
A simplified backlog-growth estimate is:
\[ BacklogGrowthRate = ProducerRate - ConsumerCompletionRate \]
When producer rate remains greater than completion rate, backlog and processing delay continue to grow.
Queue Backlog
Useful queue pressure signals include:
- Visible message count
- Oldest-message age
- In-flight message count
- Consumer completion rate
- Retry rate
- Dead-letter rate
Queue length alone can be misleading because messages can have different processing costs.
Consumer Lag
In a log, consumer lag represents the distance between the log's latest available position and the consumer group's processed or committed position.
A simplified record-based definition is:
\[ ConsumerLag = LatestLogOffset - ConsumerCommittedOffset \]
Latest partition offset:
50,000
Consumer committed offset:
46,500
Consumer lag:
3,500 records
Record count should be combined with event age and processing rate because records can differ in size and processing cost.
Drain-time Estimate
A simplified backlog-drain estimate is:
\[ EstimatedDrainTime = \frac{ Backlog }{ ConsumerRate - ProducerRate } \]
This estimate is meaningful only when consumer rate is greater than producer rate. Retries, variation in message cost, partition imbalance, and dependency limits affect actual recovery.
Queue Use Cases
| Use Case | Why a Queue Fits |
|---|---|
| Email delivery | Each delivery task needs one successful worker |
| Image resizing | Workers can compete for independent processing tasks |
| Report generation | Expensive work can be buffered and processed within concurrency limits |
| Webhook delivery | Each destination attempt can use retry and dead-letter handling |
| Batch job dispatch | Tasks can be distributed among equivalent workers |
| Media processing | Each media object requires one successful processing workflow |
Log Use Cases
| Use Case | Why a Log Fits |
|---|---|
| Domain-event distribution | Several independent services can consume the same event |
| Change Data Capture | Ordered database changes can feed several downstream systems |
| Search-index updates | The consumer can replay retained events to rebuild an index |
| Analytical ingestion | Analytics can process an independent retained event stream |
| Materialized views | A derived state can be reconstructed from retained events |
| Audit-oriented event history | Events remain available according to the defined retention policy |
Decision Matrix
| Requirement | Prefer Queue | Prefer Log |
|---|---|---|
| One successful worker per task | Strong fit | Possible through one consumer group, but may add unnecessary complexity |
| Several independent consumers need the same event | Requires routed copies or separate queues | Strong fit through separate consumer groups |
| Historical replay | Requires additional retention design | Strong fit within the retained range |
| Complex task routing | Strong fit in queue-oriented routing systems | Usually modeled with topics, keys, and consumers |
| Per-message acknowledgment workflow | Strong fit | Usually represented through offset management |
| Ordered history per entity | Requires careful queue ordering and concurrency control | Strong fit with a stable partition key |
| Rebuild a projection | Requires another source of historical messages | Replay from a retained offset |
| Task priority | Often fits separate or priority queues | Can require separate topics or application policy |
Combining Queues and Logs
A system does not need to select one messaging model for every workload.
Business transaction
|
v
Domain event appended to log
|
+-- Analytics consumer group
+-- Search consumer group
+-- Notification coordinator
|
v
Notification queue
|
+-- Email worker
+-- SMS worker
+-- Push worker
The log distributes a retained business event to independent consumers. A downstream queue then dispatches concrete tasks to competing workers.
Combination rule: Use logs to distribute retained facts to independent subscribers. Use queues to assign concrete units of work to workers.
Event vs Command
| Message Type | Meaning | Example |
|---|---|---|
| Command | A request for a specific action to be performed | GenerateCertificate |
| Event | A statement that something already happened | CertificateIssued |
Commands often fit queue semantics because one handler should perform the requested work. Events often fit log semantics because several independent consumers can react to the same fact.
Transactional Outbox
A service can update business data and record an outgoing message in the same local database transaction.
Application transaction
|
+-- Update business table
+-- Insert outbox record
|
v
Commit once
|
v
Outbox publisher reads record
|
v
Publish to queue or log
The outbox publisher can retry publication. Consumers must still handle duplicate message delivery.
Conceptual Outbox Table
CREATE TABLE message_outbox
(
message_id VARCHAR(150) PRIMARY KEY,
aggregate_type VARCHAR(100) NOT NULL,
aggregate_id VARCHAR(150) NOT NULL,
message_type VARCHAR(150) NOT NULL,
payload TEXT NOT NULL,
created_at TIMESTAMP NOT NULL,
published_at TIMESTAMP NULL
);
Example Queue Message
{
"messageId": "unique-message-id",
"commandType": "GenerateCourseCertificate",
"tenantId": 17,
"learnerId": 1042,
"courseId": 42,
"requestedAt": "message-created-time"
}
A command message should include a stable identifier, routing context, and only the data required by the consumer.
Example Log Event
{
"eventId": "unique-event-id",
"eventType": "CourseCompleted",
"eventVersion": 1,
"tenantId": 17,
"learnerId": 1042,
"courseId": 42,
"occurredAt": "event-occurrence-time"
}
An event should describe an accepted fact. Event schemas should evolve compatibly because several independently deployed consumers can read the event.
Message Schema Evolution
Producers and consumers can be deployed at different times.
Use:
- Explicit message type and version
- Backward-compatible field additions
- Stable field meanings
- Consumer tolerance for unknown optional fields
- Controlled removal of obsolete fields
- Schema validation and compatibility testing
Do not reuse an existing field with a different meaning. Introduce a new version or field instead.
Security and Privacy
Queue and log records can outlive the original request and can be consumed by several systems.
Protect messaging data by:
- Publishing only the minimum required data
- Using trusted tenant and caller context
- Encrypting transport and storage according to policy
- Applying least-privilege producer and consumer access
- Separating tenants or workloads where required
- Protecting broker-management operations
- Redacting secrets and credentials
- Defining retention and deletion policies
- Auditing access to sensitive topics and queues
Event retention must be aligned with privacy and data-lifecycle requirements. Replay capability should not become indefinite retention without a defined purpose.
Conceptual Queue Configuration
queue:
name: certificate-generation
delivery:
acknowledgment: explicit
visibilityTimeout: approved-processing-boundary
retries:
maximumAttempts: approved-limit
exponentialBackoff: true
jitter: true
deadLetter:
enabled: true
destination: certificate-generation-failures
consumers:
concurrency: approved-safe-limit
observability:
backlog: enabled
oldestMessageAge: enabled
retries: enabled
deadLetters: enabled
Conceptual Log Configuration
eventLog:
topic: learning-domain-events
partitions: approved-partition-count
replication: approved-replication-policy
producer:
key: trusted-aggregate-id
idempotence: platform-supported
retention:
policy: approved-retention-policy
consumers:
groups:
- notifications
- search-index
- analytics
observability:
partitionLag: enabled
consumerGroupLag: enabled
rebalanceEvents: enabled
retainedStorage: enabled
These configurations are conceptual. Actual ordering, retention, acknowledgment, transactional, and delivery guarantees must be verified in the selected messaging platform's documentation.
PHP Idempotent Consumer Example
<?php
declare(strict_types=1);
final class EnrollmentEventConsumer
{
public function __construct(
private DatabaseConnection $database,
private ProcessedMessageRepository $processedMessages,
private EnrollmentProjection $projection
) {
}
public function consume(
array $event
): void {
$messageId =
(string)$event['eventId'];
$this->database->transaction(
function () use (
$messageId,
$event
): void {
if (
$this->processedMessages
->exists(
consumerName:
'enrollment-projection',
messageId:
$messageId
)
) {
return;
}
$this->projection
->applyEnrollmentEvent(
$event
);
$this->processedMessages
->record(
consumerName:
'enrollment-projection',
messageId:
$messageId
);
}
);
}
}
Production code requires schema validation, secure tenant context, error classification, retry policy, database constraints, logging, and platform-specific acknowledgment or offset handling.
Learning-platform Examples
| Workload | Recommended Model | Reasoning |
|---|---|---|
| Generate a learner certificate | Queue | One worker should complete one generation task |
| Resize a course thumbnail | Queue | The work can be distributed among equivalent media workers |
| Course-completed event | Log | Notifications, analytics, certification, and recommendations can consume independently |
| Send an individual email | Queue | The delivery task needs one successful worker |
| Enrollment event history | Log | Several consumers can process retained enrollment facts |
| Rebuild the course-search index | Log | The search consumer can replay retained course events |
| Large report export | Queue | Concurrency and backlog can be bounded independently |
| Learning analytics ingestion | Log | Analytics can consume events at its own position and pace |
Observability
Queue Metrics
- Published-message rate
- Consumer completion rate
- Visible backlog
- Oldest-message age
- In-flight message count
- Redelivery rate
- Retry count
- Dead-letter count
- Processing duration
- Consumer concurrency
Log Metrics
- Records appended per partition
- Bytes appended per partition
- Consumer-group lag
- Oldest unprocessed-event age
- Partition traffic skew
- Consumer processing rate
- Offset-commit failures
- Consumer-group rebalances
- Retained storage
- Replication health
Structured Processing Event
{
"messageSystem": "event-log",
"destination": "learning-domain-events",
"consumerGroup": "search-index",
"partition": "protected-partition-reference",
"processingResult": "success",
"deliveryAttempt": 1,
"schemaVersion": 1
}
Avoid logging message payloads when identifiers and processing metadata are sufficient.
Alert Conditions
Alert when:
- Queue backlog continues growing
- Oldest queued-message age exceeds the processing objective
- Consumer completion rate falls below producer rate
- Redelivery or retry rate increases unexpectedly
- Dead-letter volume grows
- One log partition becomes hot
- Consumer-group lag continues increasing
- Consumer-group rebalances occur repeatedly
- A poison event blocks partition progress
- Required log history approaches its retention boundary
- Broker storage approaches capacity
- Replication falls below its required level
Troubleshooting Workflow
- Identify whether the destination uses queue or log semantics.
- Confirm the message or event identity.
- Check producer publication success.
- Check broker availability and storage health.
- For a queue, inspect backlog, in-flight messages, acknowledgments, and visibility.
- For a log, inspect topic, partition, latest offset, and consumer-group offset.
- Check consumer health and processing rate.
- Check retries, redeliveries, and dead-letter handling.
- Check partition or key skew.
- Check downstream database and service dependencies.
- Check idempotency and duplicate-processing records.
- Check schema compatibility.
- Check recent consumer-group or routing changes.
- Replay or redrive messages through the approved recovery procedure.
Common Messaging Mistakes
Using a Queue When Several Consumers Need the Same History
Every independent subscriber requires another routed copy or retained source.
Using a Log for a Simple One-worker Task without Need
The design inherits partition, offset, retention, and group-management complexity for basic task dispatch.
Acknowledging before Business Commit
A consumer failure can cause the broker to consider unfinished work successfully completed.
Committing an Offset before Processing Completes
A restart can skip an event whose business effect was never committed.
Assuming At-least-once Means Exactly One Effect
Redelivery can repeat the operation unless the consumer is idempotent.
Expecting Global Order across Log Partitions
Ordering is normally limited to each partition.
Choosing a Hot Partition Key
One key can overload one partition while other partitions remain underused.
Adding Consumers beyond Available Partitions
Additional consumers can remain idle because active partition ownership is already assigned.
Retrying Poison Messages Indefinitely
Repeated failure consumes capacity and can prevent useful work from progressing.
Using an Unbounded Backlog
Processing delay, storage use, and recovery duration continue increasing during sustained overload.
Keeping Events without a Retention Policy
Broker storage grows without a defined operational or business boundary.
Publishing Business Data without an Atomic Outbox
The database update can commit while publication fails, or publication can occur while the business transaction rolls back.
Recommended Test Cases
| Test | Expected Evidence |
|---|---|
| Normal queue processing | One successful worker completes each queued task |
| Worker failure before acknowledgment | The message becomes available for safe redelivery |
| Lost acknowledgment | Idempotency prevents a duplicate business effect |
| Poison message | Bounded retries move the message to the approved failure path |
| Queue overload | Backpressure and bounded intake prevent broker or consumer collapse |
| Multiple log consumer groups | Every group processes the retained event independently |
| Consumer restart | The consumer resumes from its committed position |
| Event replay | A derived projection can be rebuilt without duplicate side effects |
| Partition ordering | Events for one key are processed in their partition order |
| Hot partition | Partition skew is detected and protected |
| Consumer-group rebalance | Ownership transfers without skipping committed work |
| Schema version change | Old and new consumers process compatible messages |
| Retention boundary | Replay expectations match the events still available |
| Outbox publication failure | The committed outbox record is retried safely |
Queues and Logs Best Practices
Recommended Practices
- Use queues for work that needs one successful worker.
- Use logs for retained facts required by independent consumers.
- Model commands and events separately.
- Give every message a stable unique identifier.
- Make consumers idempotent.
- Acknowledge or commit progress only at the correct processing boundary.
- Use bounded retries with exponential backoff and jitter.
- Move poison messages to a controlled failure workflow.
- Monitor backlog age rather than message count alone.
- Monitor consumer lag by partition and consumer group.
- Select partition keys that preserve required local ordering.
- Avoid hot partition keys.
- Do not assume global ordering across partitions.
- Add partitions before existing partitions reach critical capacity.
- Keep producer and consumer schemas compatible.
- Use retention aligned with replay and data-lifecycle requirements.
- Use a transactional outbox for database changes and publication.
- Apply backpressure when consumers cannot keep pace.
- Protect brokers and downstream services with concurrency limits.
- Test redelivery, replay, rebalancing, poison messages, and recovery.
Practice Exercise
Design the messaging layer for your online learning platform.
Requirements
- List every asynchronous workflow.
- Classify each message as a command, event, or notification.
- Use a queue for certificate-generation tasks.
- Use a queue for course-thumbnail processing.
- Publish course-completed facts to a retained event log.
- Create separate notification, analytics, and search consumer groups.
- Select a partition key preserving per-learner event order.
- Add unique message and event identifiers.
- Implement idempotent consumers.
- Create bounded retries and dead-letter handling.
- Use a transactional outbox for publication.
- Monitor queue backlog and oldest-message age.
- Monitor consumer lag by group and partition.
- Test one poison message.
- Replay events to rebuild a search projection.
- Test consumer failure and group rebalancing.
Messaging-design Template
| Workflow | Message Model | Progress Tracking | Failure Protection |
|---|---|---|---|
| Generate certificate | Queue command | Acknowledgment after durable outcome | Idempotency, retry, and dead-letter handling |
| Process uploaded video | Queue command | Task completion and processing checkpoint | Bounded concurrency and retry |
| Learner enrolled | Log event | Offset per independent consumer group | Consumer idempotency and replay |
| Course published | Log event | Partition offset | Schema compatibility and retained replay |
| Send enrollment email | Queue task derived from event | Queue acknowledgment | Delivery deduplication and dead-letter workflow |
| Rebuild search index | Log replay | Dedicated consumer-group offset | Replaceable target index and checkpointed replay |
Frequently Asked Questions
What is a message queue?
A message queue stores tasks waiting for an eligible consumer, normally completing each message after one successful worker acknowledges it.
What is an event log?
An event log is an append-only sequence of retained records that consumers read using independent positions or offsets.
What is the main difference?
A queue primarily distributes tasks for completion, while a log retains events so independent consumer groups can read and replay them.
Can several workers consume one queue?
Yes. Competing consumers divide queued messages so different workers process different tasks.
Can several services consume the same log event?
Yes. Separate consumer groups can independently read the same retained event.
What is an acknowledgment?
An acknowledgment tells a queue broker that the delivered message has reached its successful processing boundary.
What is an offset?
An offset identifies a position inside a log partition and allows a consumer to track processing progress.
Can queue messages be replayed?
Not normally after successful completion unless the system retains another copy or archive designed for replay.
Can log events be replayed?
Yes, if the required events remain within the configured retention or compaction policy.
Does adding log consumers always increase throughput?
No. Within one group, active parallelism is constrained by available partitions and downstream processing capacity.
Why must consumers be idempotent?
A message or event can be delivered again after a timeout, failure, acknowledgment loss, or consumer-group reassignment.
Can queues and logs be used together?
Yes. A retained event log can distribute a business fact to independent services, and one service can create queue tasks for concrete worker actions.
Key Takeaway
Queues and logs both decouple producers from consumers, but they model different responsibilities. A queue treats a message primarily as work to be completed by one successful worker. Competing workers distribute the backlog, and acknowledgment completes the message. A log treats a message as a retained event. Independent consumer groups track their own offsets, process events at different speeds, and can replay history while the required records remain retained. Use queues for task dispatch, background jobs, email delivery, media processing, and other one-worker actions. Use logs for domain events, Change Data Capture, materialized views, search indexing, analytics, and workloads requiring independent consumption or replay. In both models, design for duplicate delivery, idempotency, backpressure, schema evolution, poison messages, security, and observability. Select the model from the message's meaning and lifecycle rather than from product popularity.