Kafka, RabbitMQ and SQS
Kafka, RabbitMQ, and Amazon SQS
Learn how Apache Kafka, RabbitMQ, and Amazon SQS support asynchronous communication using different messaging models. Compare event logs, exchanges, queues, partitions, consumer groups, acknowledgments, offsets, visibility timeouts, replay, ordering, retries, dead-letter handling, scaling, operational responsibility, and practical system-design use cases.
Introduction
Distributed applications frequently need to perform work outside the original request-response path.
Client submits request
|
v
Application validates request
|
v
Application publishes message
|
v
Application responds
|
v
Consumer processes message asynchronously
Apache Kafka, RabbitMQ, and Amazon Simple Queue Service, commonly called Amazon SQS, all help producers communicate asynchronously with consumers. However, these technologies are based on different messaging models.
- Apache Kafka is primarily a distributed event-streaming log.
- RabbitMQ is a message broker built around exchanges, queues, bindings, and acknowledgments.
- Amazon SQS is a fully managed AWS queue service offering Standard and FIFO queues.
Core idea: Choose Kafka when retained events, replay, and independent event consumers are central requirements. Choose RabbitMQ when flexible broker-side routing and traditional worker queues are central. Choose Amazon SQS when an AWS-native application needs managed queue semantics with minimal broker administration.
Prerequisites
| # | Prerequisite | Why It Is Needed |
|---|---|---|
| 1 | Queues vs logs | Kafka primarily uses retained-log semantics, while RabbitMQ and SQS commonly use queue semantics. |
| 2 | Producer-consumer pattern | Producers publish records and consumers process them. |
| 3 | Commands vs events | Commands request work, while events describe facts that already occurred. |
| 4 | Idempotency | Consumers must safely handle duplicate delivery or replay. |
| 5 | Partitioning and ordering | Kafka and SQS FIFO scope ordering using partitions or message groups. |
| 6 | Retries and dead-letter handling | Failed messages require bounded and observable recovery paths. |
| 7 | Backpressure and overload control | Message production can exceed safe consumer capacity. |
Quick Comparison
| Area | Apache Kafka | RabbitMQ | Amazon SQS |
|---|---|---|---|
| Primary model | Distributed, partitioned event log | Message broker with exchanges and queues | Fully managed cloud queue |
| Primary message meaning | Retained event or stream record | Routed message or task | Queued task or command |
| Consumption progress | Consumer-group offsets | Consumer acknowledgments | Receive, visibility timeout, and delete |
| After successful consumption | Record remains until retention or compaction removes it | Acknowledged queue message can be removed | Consumer explicitly deletes the processed message |
| Replay | Native while records remain retained | Traditional queues are completion-oriented rather than replay-oriented | Deleted messages are not available for normal replay |
| Independent subscribers | Separate consumer groups | Separate queues bound through routing rules | Separate queues, often fed through another fan-out service |
| Routing | Topics, partitions, and record keys | Exchanges, bindings, and routing keys | Queue destination, with Standard or FIFO behaviour |
| Ordering scope | Within a partition | Depends on queue, delivery, concurrency, retry, and configuration | Best effort for Standard; FIFO ordering through FIFO queue semantics and message groups |
| Operations | Requires Kafka infrastructure or a managed Kafka offering | Requires RabbitMQ infrastructure or a managed RabbitMQ offering | AWS manages the queue service |
| Strongest fit | Event streaming, replay, CDC, analytics, and multiple independent consumers | Task queues, flexible routing, and worker distribution | AWS-native background processing and service decoupling |
Apache Kafka
Apache Kafka organizes records into topics. Each topic is divided into one or more partitions, and each partition is an ordered append-only log.
Kafka topic:
learning-domain-events
Partition 0:
Offset 0 -> CourseCreated
Offset 1 -> CoursePublished
Offset 2 -> CourseArchived
Partition 1:
Offset 0 -> LearnerEnrolled
Offset 1 -> LessonCompleted
Offset 2 -> CourseCompleted
Consumers read records using offsets. Reading a record does not remove it from the topic.
Main Kafka Components
| Component | Purpose |
|---|---|
| Producer | Publishes records to Kafka topics |
| Broker | Stores topic partitions and serves producer and consumer requests |
| Topic | Logical event stream containing related records |
| Partition | Ordered unit of storage and consumer parallelism |
| Offset | Position of a record inside a partition |
| Consumer group | Consumers cooperating to divide partition processing |
| Kafka Connect | Framework that connects Kafka with supported external systems through connectors |
Kafka Write Path
Producer creates event
|
v
Select topic
|
v
Select partition using key
or partitioning policy
|
v
Kafka appends record
to partition log
|
v
Record is replicated
according to cluster policy
|
v
Producer receives result
according to acknowledgment policy
Kafka Read Path
Consumer joins group
|
v
Partitions are assigned
|
v
Consumer fetches records
from current offsets
|
v
Consumer processes records
|
v
Consumer commits progress
Kafka Consumer Groups
Consumers inside one group divide topic partitions. Separate groups consume the same retained stream independently.
Topic:
Partition 0
Partition 1
Partition 2
Partition 3
Notification Group:
Consumer N1 -> Partitions 0 and 1
Consumer N2 -> Partitions 2 and 3
Analytics Group:
Consumer A1 -> Partition 0
Consumer A2 -> Partition 1
Consumer A3 -> Partition 2
Consumer A4 -> Partition 3
A partition is assigned to one active consumer within a group at a given time. Adding consumers beyond the group's available partitions does not create additional active partition ownership.
Kafka Offsets and Replay
Kafka consumers commit offsets so that processing can resume after restart or reassignment.
Partition latest offset:
25,000
Committed group offset:
23,500
Consumer lag:
1,500 records
A consumer can replay retained records by beginning from an earlier valid offset.
Current offset:
25,000
Replay offset:
20,000
Consumer reprocesses:
20,000 through 25,000
Replay is useful for:
- Rebuilding a search index
- Recreating a materialized view
- Correcting a previously defective consumer
- Feeding a new downstream system
- Reprocessing analytical data
Kafka rule: Replay is possible only while the required topic records and consumer-position information remain available. Retention and consumer recovery policy must match the business recovery requirement.
Kafka Ordering
Kafka provides ordering inside an individual partition.
Partition for learner-1042:
Offset 100 -> CourseStarted
Offset 101 -> LessonCompleted
Offset 102 -> QuizPassed
Offset 103 -> CourseCompleted
Kafka does not automatically provide one global order across all partitions. Events requiring local ordering should use a stable key that places them in the same partition.
Kafka Hot Partitions
A low-cardinality or highly popular record key can overload one partition.
Normal keys:
1,000 events per minute
One popular key:
500,000 events per minute
Result:
The partition owning the key
becomes the bottleneck.
Possible controls include:
- Choose a higher-cardinality partition key
- Split safely independent work into subkeys
- Create a controlled bucket component
- Separate very heavy workloads into another topic
- Apply producer rate limits
- Add partitions before current partition capacity is exhausted
Kafka Connect
Kafka Connect provides a framework for transferring data between Kafka and supported external systems using source and sink connectors.
Source system
|
v
Source connector
|
v
Kafka topic
|
v
Sink connector
|
v
Target system
Source connectors can bring records into Kafka. Sink connectors can consume Kafka records and transfer them to external systems such as databases, search indexes, files, or HTTP endpoints, depending on the connector.
Kafka Strengths
- Retained and replayable event streams
- Independent consumer groups
- Partition-based parallelism
- Ordering within a partition
- Strong fit for Change Data Capture and stream processing
- Suitable foundation for multiple real-time downstream pipelines
- Connector ecosystem for supported source and target systems
Kafka Trade-offs
- Topics, partitions, retention, replication, and consumer groups require operational design
- Consumer-group rebalancing must be handled correctly
- Partition-count and partition-key choices affect throughput and ordering
- Hot partitions can limit parallelism
- Kafka is usually more complex than a simple managed task queue
- Task priority and complex per-message routing can require additional design
RabbitMQ
RabbitMQ is a message broker in which producers commonly publish to an exchange. The exchange routes messages to one or more queues using its type, bindings, and routing information.
Producer
|
v
Exchange
|
+-- Queue A -> Consumer Group A
+-- Queue B -> Consumer Group B
+-- Queue C -> Consumer Group C
Main RabbitMQ Components
| Component | Purpose |
|---|---|
| Producer | Publishes a message to an exchange |
| Exchange | Applies routing rules to published messages |
| Binding | Connects an exchange to a queue, stream, or another exchange using routing rules |
| Queue | Stores messages waiting for consumers |
| Consumer | Receives and processes messages from a queue |
| Acknowledgment | Confirms that a consumer successfully processed a delivery |
| Publisher confirm | Confirms the publisher's interaction with the RabbitMQ node and relevant queue or stream leader |
RabbitMQ Exchange Types
| Exchange Type | General Routing Behaviour |
|---|---|
| Direct | Routes according to an exact routing-key and binding-key match |
| Topic | Routes using pattern matching over routing keys |
| Fanout | Routes a copy to every bound destination |
| Headers | Routes according to configured message-header matching |
Topic-routing Example
Message routing key:
learning.course.published
Binding:
learning.course.*
Result:
Message matches the binding
and is routed to the queue.
RabbitMQ Acknowledgments
Consumer acknowledgments tell RabbitMQ when a delivery has been processed successfully.
RabbitMQ delivers message
|
v
Consumer processes message
|
+-- Success:
| send acknowledgment
| message can be removed
|
+-- Failure:
reject or negatively acknowledge
according to retry policy
Publisher confirms and consumer acknowledgments solve different parts of the delivery path.
- Publisher confirms address publication from publisher to RabbitMQ.
- Consumer acknowledgments address processing from RabbitMQ to consumer.
RabbitMQ rule: Confirming publication does not prove that a consumer completed the business operation. Publisher confirms and consumer acknowledgments are separate mechanisms.
RabbitMQ Prefetch
Prefetch limits the number of unacknowledged deliveries sent to a consumer or channel, according to the client and broker configuration.
Consumer prefetch:
10
Maximum assigned but
unacknowledged deliveries:
10
A very large prefetch can overload a slow consumer and distribute work unevenly. A very small prefetch can reduce throughput for fast consumers. Tune it using measured processing duration and memory consumption.
RabbitMQ Competing Consumers
Task Queue
|
+-- Worker A
+-- Worker B
+-- Worker C
Consumers connected to one queue compete for its messages. If several independent services need the same published message, the exchange can route copies to separate queues.
RabbitMQ Strengths
- Flexible exchange and binding model
- Strong fit for task and worker queues
- Competing-consumer processing
- Publisher confirms and consumer acknowledgments
- Suitable for direct, topic, fanout, and header-based routing
- Useful when the broker should make routing decisions
RabbitMQ Trade-offs
- Traditional queues are not primarily designed as retained event history
- Each independent subscriber commonly needs its own queue
- Broker topology and queue lifecycle require active management
- Incorrect acknowledgment or requeue behaviour can cause loss or repeated delivery
- Poison messages can create repeated redelivery
- Applications must tune prefetch and consumer concurrency
Amazon SQS
Amazon SQS is a managed AWS queue service. Producers send messages to a queue, and consumers poll the queue for available messages.
Producer
|
v
Amazon SQS Queue
|
+-- Consumer A
+-- Consumer B
+-- Consumer C
Amazon SQS provides two queue types:
- Standard queues
- FIFO queues
SQS Standard Queues
Standard queues are designed for very high throughput, at-least-once delivery, and best-effort ordering.
Applications using Standard queues must handle:
- Occasional duplicate delivery
- Messages arriving in a different order
- Idempotent processing
- Visibility timeout expiration
- Explicit message deletion after success
SQS FIFO Queues
FIFO queues are designed for workloads where ordering and duplicate suppression within the queue's documented processing model are important.
FIFO queues use message groups to provide separate ordered sequences inside one queue.
FIFO Queue
Message Group: learner-1042
RegisterAccount
EnrollInCourse
CompleteCourse
Message Group: learner-2055
RegisterAccount
EnrollInCourse
CompleteCourse
Messages inside one group are processed in order. Different message groups can support parallel progress.
SQS Standard vs FIFO
| Area | Standard Queue | FIFO Queue |
|---|---|---|
| Ordering | Best effort | FIFO semantics within the applicable message-group scope |
| Delivery | At least once; duplicates can occur | Designed to prevent duplicate messages from being introduced into the queue under its deduplication model |
| Parallelism | Consumers process available messages broadly | Parallelism comes from several independent message groups |
| Primary fit | High-volume work that tolerates duplicates and reordering | Work requiring ordered processing within a business key |
SQS Visibility Timeout
When a consumer receives an SQS message, the message becomes temporarily invisible to other consumers.
Consumer receives message
|
v
Visibility timeout begins
|
+-- Processing succeeds:
| consumer deletes message
|
+-- Message is not deleted:
visibility timeout expires
message becomes available again
If the visibility timeout is shorter than normal processing duration, another consumer can receive the same message before the first consumer finishes.
If the visibility timeout is unnecessarily long, failed work can take longer to become eligible for retry.
SQS Delete after Success
Receiving a message does not mean that processing has completed. The consumer should delete the message only after the required business outcome has completed successfully.
Receive message
|
v
Validate payload
|
v
Process idempotently
|
v
Commit business outcome
|
v
Delete message
SQS Dead-letter Queues
An SQS queue can use a dead-letter queue for messages that are not processed successfully after the configured receive attempts.
Source Queue
|
v
Consumer processing fails
|
v
Message becomes visible again
|
v
Receive count reaches policy
|
v
Move to Dead-letter Queue
A dead-letter workflow should include monitoring, investigation, correction, and controlled redrive.
SQS Strengths
- Managed AWS queue infrastructure
- Standard and FIFO queue options
- Visibility timeout for in-progress work
- Dead-letter queue support
- Strong fit for AWS-native background processing
- Consumers can scale independently from producers
- No messaging-broker cluster administration by the application team
SQS Trade-offs
- It is a queue service rather than a replayable event-streaming log
- Consumers poll queues rather than reading retained partition logs
- Standard messages can be duplicated or delivered out of order
- FIFO parallelism depends on appropriate message-group design
- Complex fanout and routing can require additional AWS services or queues
- Applications must tune visibility, retries, deletion, and dead-letter behaviour
Detailed Feature Comparison
| Capability | Kafka | RabbitMQ | SQS |
|---|---|---|---|
| Event replay | Core capability through retained records and offsets | Not the primary model of traditional queues | Deleted queue messages are not replayed normally |
| Independent subscribers | Separate consumer groups | Separate routed queues | Separate queues |
| Task distribution | Possible using one consumer group | Core competing-consumer pattern | Core queue pattern |
| Broker-side routing | Topics, keys, and partitions | Rich exchanges and bindings | Queue-level destination |
| Processing completion | Consumer commits offset according to processing design | Consumer acknowledgment | Consumer deletes message |
| Temporary processing lock | Partition assignment and consumer-group ownership | Unacknowledged delivery state | Visibility timeout |
| Ordering | Within a partition | Depends on queue topology and processing model | Best effort in Standard; FIFO behaviour in FIFO queues |
| Failure isolation | Topics, partitions, and consumer groups | Virtual hosts, exchanges, queues, and consumers | Separate queues and dead-letter queues |
| Managed infrastructure | Available through managed Kafka products | Available through managed RabbitMQ products | Native characteristic of the AWS service |
Choosing the Correct Platform
Choose Kafka When
- Events must remain replayable
- Several independent consumer groups require the same event stream
- Change Data Capture is a major requirement
- Real-time stream processing is required
- Ordering is needed within a stable business key
- Historical events must rebuild projections or indexes
- Event-stream retention is part of the architecture
Choose RabbitMQ When
- Messages require flexible exchange-based routing
- Commands must be distributed among competing workers
- Direct, topic, fanout, or header-based routing is useful
- Per-message acknowledgment and requeue behaviour are important
- The team can operate RabbitMQ or use an appropriate managed offering
- Retained event replay is not the central requirement
Choose Amazon SQS When
- The solution runs primarily on AWS
- A managed queue is preferred over operating brokers
- The workload is background-task processing
- Producer and consumer scaling should remain independent
- Standard queue duplicates and reordering can be handled
- FIFO message groups meet the required ordering scope
- Visibility timeout and dead-letter queue semantics fit the workflow
Selection rule: Select the platform from message meaning, replay requirements, routing complexity, ordering scope, consumer model, cloud environment, and operational responsibility. Do not select it only from benchmark numbers or product popularity.
Using More than One Platform
One organization can use different messaging technologies for different workloads.
Business transaction
|
v
Kafka event:
CourseCompleted
|
+-- Analytics consumer group
+-- Search consumer group
+-- Certification service
|
v
RabbitMQ or SQS task:
GenerateCertificate
|
v
Competing workers
Kafka distributes and retains the business fact. RabbitMQ or SQS can then dispatch a concrete unit of work.
Command and Event Examples
| Message | Type | Possible Platform Direction |
|---|---|---|
CourseCompleted |
Event | Kafka when several independent consumers need the retained fact |
GenerateCertificate |
Command | RabbitMQ or SQS when one worker should perform the task |
SendEnrollmentEmail |
Command | RabbitMQ or SQS worker queue |
LearnerEnrolled |
Event | Kafka when analytics, notifications, and search consume independently |
ResizeCourseImage |
Command | RabbitMQ or SQS task queue |
Delivery and Idempotency
All three systems require application-level attention to duplicate processing and uncertain outcomes.
Consumer performs business update
|
v
Update commits
|
v
Acknowledgment, offset commit,
or delete request fails
|
v
Message can be processed again
Use:
- Stable message or event identifiers
- Unique database constraints
- Conditional writes
- Processed-message records
- Idempotency keys
- Transactional consumer updates where supported
Idempotent Message
{
"messageId": "stable-unique-message-id",
"messageType": "GenerateCertificate",
"messageVersion": 1,
"tenantId": 17,
"learnerId": 1042,
"courseId": 42
}
Transactional Outbox
A transactional outbox helps prevent a service from committing business data without recording the corresponding message.
Local database transaction
|
+-- Update business record
+-- Insert outbox message
|
v
Commit once
|
v
Publisher reads outbox
|
v
Publish to Kafka,
RabbitMQ, or SQS
|
v
Mark publication progress
The publisher can retry after failure. Consumers must still be idempotent because publication can occur more than once.
Backpressure
A messaging platform buffers work, but buffering does not create unlimited consumer capacity.
Let \(P\) be the production rate and \(C\) be the successful consumer completion rate.
Backlog grows when:
\[ P > C \]
A simplified backlog growth rate is:
\[ BacklogGrowthRate = P - C \]
Use:
- Producer rate limits
- Bounded consumer concurrency
- Autoscaling based on backlog or lag
- Batching where supported and appropriate
- Load shedding for optional work
- Dead-letter handling for poison messages
Conceptual Kafka Configuration
kafka:
topic: learning-domain-events
partitions: approved-partition-count
replication: approved-replication-policy
producer:
key: trusted-aggregate-id
acknowledgments: approved-durability-policy
consumer:
groupId: analytics-projection
offsetCommit: after-durable-processing
retention:
policy: approved-replay-requirement
observability:
consumerLag: enabled
partitionSkew: enabled
rebalanceEvents: enabled
retainedStorage: enabled
Conceptual RabbitMQ Configuration
rabbitmq:
exchange:
name: learning.commands
type: topic
queue:
name: certificate-generation
durable: true
binding:
routingKey: certificate.generate
publisher:
confirms: enabled
consumer:
acknowledgment: manual
prefetch: approved-safe-limit
failureHandling:
boundedRetries: true
deadLetterDestination: certificate-failures
observability:
readyMessages: enabled
unacknowledgedMessages: enabled
redeliveries: enabled
Conceptual SQS Configuration
sqs:
queueType: standard-or-fifo
queueName: certificate-generation
processing:
visibilityTimeout: approved-processing-boundary
deleteAfterDurableSuccess: true
idempotentConsumerRequired: true
fifo:
messageGroupKey: trusted-business-key
deduplicationKey: stable-message-id
failureHandling:
deadLetterQueue: certificate-generation-failures
maximumReceives: approved-limit
observability:
availableMessages: enabled
inFlightMessages: enabled
oldestMessageAge: enabled
deadLetterMessages: enabled
These are conceptual examples. Exact settings and supported guarantees must be verified against the deployed platform and version.
PHP Idempotent Consumer
<?php
declare(strict_types=1);
final class CertificateCommandHandler
{
public function __construct(
private DatabaseConnection $database,
private ProcessedMessageRepository $processedMessages,
private CertificateService $certificates
) {
}
public function handle(
array $message
): void {
$messageId =
(string)$message['messageId'];
$this->database->transaction(
function () use (
$messageId,
$message
): void {
if (
$this->processedMessages
->exists(
consumerName:
'certificate-command-handler',
messageId:
$messageId
)
) {
return;
}
$this->certificates
->generateIfNotExists(
tenantId:
(int)$message['tenantId'],
learnerId:
(int)$message['learnerId'],
courseId:
(int)$message['courseId']
);
$this->processedMessages
->record(
consumerName:
'certificate-command-handler',
messageId:
$messageId
);
}
);
}
}
Broker acknowledgment, offset commit, or SQS deletion should happen only after this durable processing boundary succeeds.
Security Considerations
- Grant producers access only to required topics, exchanges, or queues.
- Grant consumers access only to their required destinations.
- Encrypt messaging traffic and stored messages according to policy.
- Do not publish passwords, access tokens, or reusable credentials.
- Use trusted tenant context rather than client-supplied routing identity.
- Protect broker-management and queue-management operations.
- Define event retention according to privacy and business requirements.
- Redact sensitive payloads from processing logs and dead-letter diagnostics.
- Audit administrative changes and message-access permissions.
Observability
Kafka Metrics
- Records produced by topic and partition
- Consumer-group lag
- Oldest unprocessed-event age
- Partition traffic skew
- Consumer processing rate
- Offset-commit failures
- Consumer-group rebalances
- Broker storage and replication health
RabbitMQ Metrics
- Ready-message count
- Unacknowledged-message count
- Publish and delivery rates
- Consumer acknowledgment rate
- Redelivery rate
- Queue depth and message age
- Dead-letter volume
- Consumer and connection health
SQS Metrics
- Available-message count
- In-flight-message count
- Age of oldest message
- Messages sent, received, and deleted
- Receive count and retry behaviour
- Dead-letter queue depth
- Consumer completion rate
- Visibility-timeout expirations inferred from redelivery
Alert Conditions
Alert when:
- Kafka consumer lag continues increasing
- One Kafka partition receives disproportionate traffic
- Kafka consumers rebalance repeatedly
- RabbitMQ ready or unacknowledged messages continue growing
- RabbitMQ redelivery volume increases unexpectedly
- SQS oldest-message age exceeds the processing objective
- SQS messages repeatedly return after visibility timeout
- A dead-letter queue receives unexpected messages
- Consumer completion rate remains below producer rate
- Poison messages block useful processing
- Messaging storage or broker resources approach capacity
- Schema compatibility errors increase
Common Selection Mistakes
Choosing Kafka for Every Background Task
A simple worker queue can be easier when replay and independent event consumers are not required.
Choosing RabbitMQ for Long-term Event Replay
Traditional acknowledged queues are completion-oriented rather than retained-history oriented.
Using SQS Standard while Assuming Strict Order
Standard queues use best-effort ordering and can deliver a message more than once.
Using One SQS FIFO Message Group for Every Message
One group serializes ordered processing and can limit available parallelism.
Assuming Kafka Provides Global Topic Order
Kafka ordering is scoped to each partition.
Adding Kafka Consumers beyond Partition Count
Additional consumers in the same group can remain without an active partition assignment.
Confusing RabbitMQ Publisher Confirms with Consumer Completion
Publisher confirms do not prove that a consumer completed the business operation.
Using Automatic Acknowledgment before RabbitMQ Processing
A consumer crash can lose work that RabbitMQ already considers delivered.
Deleting an SQS Message before Business Commit
The task can disappear before its intended outcome is durable.
Ignoring Idempotency
Redelivery, replay, acknowledgment loss, or offset recovery can repeat a business operation.
Using Unbounded Retries
Poison messages consume capacity and prevent useful work from progressing.
Selecting from Throughput Claims Alone
Replay, routing, ordering, management, cost, and failure recovery can be more important than raw message rate.
Recommended Test Cases
| Test | Expected Evidence |
|---|---|
| Kafka consumer restart | The consumer resumes according to its committed offset |
| Kafka replay | A projection can be rebuilt without duplicate external effects |
| Kafka hot partition | Partition skew is detected and protected |
| Kafka group rebalance | Partition ownership changes without silently skipping required records |
| RabbitMQ consumer failure | An unacknowledged message is requeued or redelivered according to policy |
| RabbitMQ lost acknowledgment | Idempotency prevents a duplicate business effect |
| RabbitMQ routing | Direct, topic, or fanout rules deliver messages only to intended queues |
| RabbitMQ poison message | Bounded attempts move the message to the approved failure path |
| SQS visibility timeout | A failed consumer causes the message to become available again |
| SQS duplicate delivery | Consumer idempotency prevents a duplicate outcome |
| SQS Standard ordering | The application remains correct when messages arrive out of order |
| SQS FIFO groups | Messages remain ordered within each group while separate groups progress independently |
| Dead-letter workflow | Failed messages are observable, diagnosable, and safely redriven |
| Producer surge | Backlog controls preserve broker, consumer, and downstream stability |
| Schema evolution | Old and new consumers remain compatible during deployment |
Best Practices
Recommended Practices
- Choose the platform from the messaging model, not popularity.
- Use Kafka for retained, replayable event streams.
- Use RabbitMQ for flexible routing and competing task consumers.
- Use SQS for managed AWS-native queues.
- Separate commands from events.
- Give every message a stable unique identifier.
- Make all consumers idempotent.
- Commit offsets, acknowledge deliveries, or delete messages only after durable processing.
- Use a stable Kafka partition key for required per-entity ordering.
- Avoid Kafka hot partitions.
- Tune RabbitMQ prefetch from measured consumer behaviour.
- Use RabbitMQ publisher confirms and consumer acknowledgments for their separate purposes.
- Set SQS visibility according to processing requirements.
- Use several SQS FIFO message groups when ordered streams can progress independently.
- Use bounded retries with exponential backoff and jitter.
- Use dead-letter handling for poison messages.
- Use the transactional outbox for reliable publication.
- Version message and event schemas.
- Monitor backlog age and consumer lag, not count alone.
- Test replay, redelivery, rebalancing, visibility expiry, and recovery.
Practice Exercise
Select Kafka, RabbitMQ, or Amazon SQS for the asynchronous workflows in your online learning platform.
Requirements
- Publish learner-domain events to Kafka.
- Create independent event consumers for analytics and search.
- Use learner ID as the Kafka key where per-learner ordering is required.
- Replay retained events to rebuild a search index.
- Create a RabbitMQ topic exchange for complex notification routing.
- Create separate RabbitMQ queues for email and push notifications.
- Configure manual acknowledgments and bounded prefetch.
- Use SQS for AWS-native course-report generation.
- Compare SQS Standard and FIFO for the report workflow.
- Create an SQS dead-letter queue.
- Add idempotency to every consumer.
- Use a transactional outbox for event and command publication.
- Test poison-message handling on all selected platforms.
- Monitor Kafka lag, RabbitMQ queue depth, and SQS oldest-message age.
- Document why each workflow uses its selected platform.
Selection Template
| Workflow | Suggested Platform Direction | Reason |
|---|---|---|
| Domain-event distribution | Kafka | Independent consumers and replayable history |
| Search-index update stream | Kafka | Consumer can replay after index recreation |
| Complex topic-based notification routing | RabbitMQ | Exchange and binding model supports flexible routing |
| Competing media-processing workers | RabbitMQ or SQS | Each task requires one successful worker |
| AWS-native serverless background job | SQS | Managed AWS queue with worker decoupling |
| Ordered AWS task stream per learner | SQS FIFO | Message groups can preserve per-learner ordering |
Frequently Asked Questions
Is Kafka a traditional message queue?
Kafka is primarily a distributed event-streaming log. It can distribute work through consumer groups, but records remain according to retention rather than being removed when one consumer reads them.
What is RabbitMQ best suited for?
RabbitMQ is well suited to traditional task queues and routing messages through exchanges and bindings to one or more queues.
What is Amazon SQS best suited for?
Amazon SQS is well suited to AWS-native applications requiring a managed queue for asynchronous task processing and service decoupling.
Which platform supports replay most naturally?
Kafka supports replay through retained topic records and consumer offsets.
Which platform has the richest broker-side routing?
RabbitMQ provides exchanges and bindings supporting direct, topic, fanout, header-based, and additional routing behaviours.
What is an SQS visibility timeout?
It is the period during which a received message remains hidden from other consumers while one consumer processes it.
Does Kafka guarantee global ordering?
No. Ordering is provided within each partition rather than across an entire multi-partition topic.
Does SQS Standard guarantee ordering?
No. Standard queues provide best-effort ordering and applications must tolerate messages arriving out of order.
How does RabbitMQ know that processing succeeded?
A consumer sends an acknowledgment after reaching its successful processing boundary.
Why do consumers still need idempotency?
A failure after the business operation but before offset commit, acknowledgment, or message deletion can cause processing to occur again.
Can these platforms be used together?
Yes. Kafka can distribute retained domain events, while RabbitMQ or SQS dispatches concrete tasks derived from those events.
Which one should I choose?
Choose Kafka for retained event streaming and replay, RabbitMQ for flexible broker routing and task queues, or SQS for managed AWS-native queuing. Validate the final choice against ordering, delivery, throughput, security, cost, and operational requirements.
Key Takeaway
Kafka, RabbitMQ, and Amazon SQS all support asynchronous communication, but they use different models. Kafka stores retained records in partitioned topics and lets independent consumer groups track offsets and replay events. RabbitMQ routes producer messages through exchanges and bindings into queues, where consumers acknowledge completed deliveries. Amazon SQS provides managed Standard and FIFO queues, using visibility timeout and explicit deletion to control processing. Choose Kafka for event streams, Change Data Capture, replay, analytics, and multiple independent subscribers. Choose RabbitMQ for flexible routing and competing worker queues. Choose SQS for AWS-native task processing without operating broker infrastructure. Regardless of platform, use idempotent consumers, transactional publication, bounded retries, dead-letter handling, schema versioning, backpressure, least-privilege access, and workload-specific observability.