Table of Contents

    Kafka, RabbitMQ and SQS

    MESSAGING & ASYNCHRONOUS PROCESSING

    Kafka, RabbitMQ, and Amazon SQS

    Learn how Apache Kafka, RabbitMQ, and Amazon SQS support asynchronous communication using different messaging models. Compare event logs, exchanges, queues, partitions, consumer groups, acknowledgments, offsets, visibility timeouts, replay, ordering, retries, dead-letter handling, scaling, operational responsibility, and practical system-design use cases.

    Introduction

    Distributed applications frequently need to perform work outside the original request-response path.

    Client submits request
          |
          v
    Application validates request
          |
          v
    Application publishes message
          |
          v
    Application responds
          |
          v
    Consumer processes message asynchronously

    Apache Kafka, RabbitMQ, and Amazon Simple Queue Service, commonly called Amazon SQS, all help producers communicate asynchronously with consumers. However, these technologies are based on different messaging models.

    • Apache Kafka is primarily a distributed event-streaming log.
    • RabbitMQ is a message broker built around exchanges, queues, bindings, and acknowledgments.
    • Amazon SQS is a fully managed AWS queue service offering Standard and FIFO queues.

    Core idea: Choose Kafka when retained events, replay, and independent event consumers are central requirements. Choose RabbitMQ when flexible broker-side routing and traditional worker queues are central. Choose Amazon SQS when an AWS-native application needs managed queue semantics with minimal broker administration.

    Prerequisites

    # Prerequisite Why It Is Needed
    1 Queues vs logs Kafka primarily uses retained-log semantics, while RabbitMQ and SQS commonly use queue semantics.
    2 Producer-consumer pattern Producers publish records and consumers process them.
    3 Commands vs events Commands request work, while events describe facts that already occurred.
    4 Idempotency Consumers must safely handle duplicate delivery or replay.
    5 Partitioning and ordering Kafka and SQS FIFO scope ordering using partitions or message groups.
    6 Retries and dead-letter handling Failed messages require bounded and observable recovery paths.
    7 Backpressure and overload control Message production can exceed safe consumer capacity.

    Quick Comparison

    Area Apache Kafka RabbitMQ Amazon SQS
    Primary model Distributed, partitioned event log Message broker with exchanges and queues Fully managed cloud queue
    Primary message meaning Retained event or stream record Routed message or task Queued task or command
    Consumption progress Consumer-group offsets Consumer acknowledgments Receive, visibility timeout, and delete
    After successful consumption Record remains until retention or compaction removes it Acknowledged queue message can be removed Consumer explicitly deletes the processed message
    Replay Native while records remain retained Traditional queues are completion-oriented rather than replay-oriented Deleted messages are not available for normal replay
    Independent subscribers Separate consumer groups Separate queues bound through routing rules Separate queues, often fed through another fan-out service
    Routing Topics, partitions, and record keys Exchanges, bindings, and routing keys Queue destination, with Standard or FIFO behaviour
    Ordering scope Within a partition Depends on queue, delivery, concurrency, retry, and configuration Best effort for Standard; FIFO ordering through FIFO queue semantics and message groups
    Operations Requires Kafka infrastructure or a managed Kafka offering Requires RabbitMQ infrastructure or a managed RabbitMQ offering AWS manages the queue service
    Strongest fit Event streaming, replay, CDC, analytics, and multiple independent consumers Task queues, flexible routing, and worker distribution AWS-native background processing and service decoupling

    Apache Kafka

    Apache Kafka organizes records into topics. Each topic is divided into one or more partitions, and each partition is an ordered append-only log.

    Kafka topic:
    
    learning-domain-events
    
    
    Partition 0:
    
    Offset 0 -> CourseCreated
    Offset 1 -> CoursePublished
    Offset 2 -> CourseArchived
    
    
    Partition 1:
    
    Offset 0 -> LearnerEnrolled
    Offset 1 -> LessonCompleted
    Offset 2 -> CourseCompleted

    Consumers read records using offsets. Reading a record does not remove it from the topic.

    Main Kafka Components

    Component Purpose
    Producer Publishes records to Kafka topics
    Broker Stores topic partitions and serves producer and consumer requests
    Topic Logical event stream containing related records
    Partition Ordered unit of storage and consumer parallelism
    Offset Position of a record inside a partition
    Consumer group Consumers cooperating to divide partition processing
    Kafka Connect Framework that connects Kafka with supported external systems through connectors

    Kafka Write Path

    Producer creates event
          |
          v
    Select topic
          |
          v
    Select partition using key
    or partitioning policy
          |
          v
    Kafka appends record
    to partition log
          |
          v
    Record is replicated
    according to cluster policy
          |
          v
    Producer receives result
    according to acknowledgment policy

    Kafka Read Path

    Consumer joins group
          |
          v
    Partitions are assigned
          |
          v
    Consumer fetches records
    from current offsets
          |
          v
    Consumer processes records
          |
          v
    Consumer commits progress

    Kafka Consumer Groups

    Consumers inside one group divide topic partitions. Separate groups consume the same retained stream independently.

    Topic:
    
    Partition 0
    Partition 1
    Partition 2
    Partition 3
    
    
    Notification Group:
    
    Consumer N1 -> Partitions 0 and 1
    Consumer N2 -> Partitions 2 and 3
    
    
    Analytics Group:
    
    Consumer A1 -> Partition 0
    Consumer A2 -> Partition 1
    Consumer A3 -> Partition 2
    Consumer A4 -> Partition 3

    A partition is assigned to one active consumer within a group at a given time. Adding consumers beyond the group's available partitions does not create additional active partition ownership.

    Kafka Offsets and Replay

    Kafka consumers commit offsets so that processing can resume after restart or reassignment.

    Partition latest offset:
    
    25,000
    
    
    Committed group offset:
    
    23,500
    
    
    Consumer lag:
    
    1,500 records

    A consumer can replay retained records by beginning from an earlier valid offset.

    Current offset:
    
    25,000
    
    
    Replay offset:
    
    20,000
    
    
    Consumer reprocesses:
    
    20,000 through 25,000

    Replay is useful for:

    • Rebuilding a search index
    • Recreating a materialized view
    • Correcting a previously defective consumer
    • Feeding a new downstream system
    • Reprocessing analytical data

    Kafka rule: Replay is possible only while the required topic records and consumer-position information remain available. Retention and consumer recovery policy must match the business recovery requirement.

    Kafka Ordering

    Kafka provides ordering inside an individual partition.

    Partition for learner-1042:
    
    Offset 100 -> CourseStarted
    Offset 101 -> LessonCompleted
    Offset 102 -> QuizPassed
    Offset 103 -> CourseCompleted

    Kafka does not automatically provide one global order across all partitions. Events requiring local ordering should use a stable key that places them in the same partition.

    Kafka Hot Partitions

    A low-cardinality or highly popular record key can overload one partition.

    Normal keys:
    
    1,000 events per minute
    
    
    One popular key:
    
    500,000 events per minute
    
    
    Result:
    
    The partition owning the key
    becomes the bottleneck.

    Possible controls include:

    • Choose a higher-cardinality partition key
    • Split safely independent work into subkeys
    • Create a controlled bucket component
    • Separate very heavy workloads into another topic
    • Apply producer rate limits
    • Add partitions before current partition capacity is exhausted

    Kafka Connect

    Kafka Connect provides a framework for transferring data between Kafka and supported external systems using source and sink connectors.

    Source system
          |
          v
    Source connector
          |
          v
    Kafka topic
          |
          v
    Sink connector
          |
          v
    Target system

    Source connectors can bring records into Kafka. Sink connectors can consume Kafka records and transfer them to external systems such as databases, search indexes, files, or HTTP endpoints, depending on the connector.

    Kafka Strengths

    • Retained and replayable event streams
    • Independent consumer groups
    • Partition-based parallelism
    • Ordering within a partition
    • Strong fit for Change Data Capture and stream processing
    • Suitable foundation for multiple real-time downstream pipelines
    • Connector ecosystem for supported source and target systems

    Kafka Trade-offs

    • Topics, partitions, retention, replication, and consumer groups require operational design
    • Consumer-group rebalancing must be handled correctly
    • Partition-count and partition-key choices affect throughput and ordering
    • Hot partitions can limit parallelism
    • Kafka is usually more complex than a simple managed task queue
    • Task priority and complex per-message routing can require additional design

    RabbitMQ

    RabbitMQ is a message broker in which producers commonly publish to an exchange. The exchange routes messages to one or more queues using its type, bindings, and routing information.

    Producer
        |
        v
    Exchange
        |
        +-- Queue A -> Consumer Group A
        +-- Queue B -> Consumer Group B
        +-- Queue C -> Consumer Group C

    Main RabbitMQ Components

    Component Purpose
    Producer Publishes a message to an exchange
    Exchange Applies routing rules to published messages
    Binding Connects an exchange to a queue, stream, or another exchange using routing rules
    Queue Stores messages waiting for consumers
    Consumer Receives and processes messages from a queue
    Acknowledgment Confirms that a consumer successfully processed a delivery
    Publisher confirm Confirms the publisher's interaction with the RabbitMQ node and relevant queue or stream leader

    RabbitMQ Exchange Types

    Exchange Type General Routing Behaviour
    Direct Routes according to an exact routing-key and binding-key match
    Topic Routes using pattern matching over routing keys
    Fanout Routes a copy to every bound destination
    Headers Routes according to configured message-header matching

    Topic-routing Example

    Message routing key:
    
    learning.course.published
    
    
    Binding:
    
    learning.course.*
    
    
    Result:
    
    Message matches the binding
    and is routed to the queue.

    RabbitMQ Acknowledgments

    Consumer acknowledgments tell RabbitMQ when a delivery has been processed successfully.

    RabbitMQ delivers message
          |
          v
    Consumer processes message
          |
          +-- Success:
          |      send acknowledgment
          |      message can be removed
          |
          +-- Failure:
                 reject or negatively acknowledge
                 according to retry policy

    Publisher confirms and consumer acknowledgments solve different parts of the delivery path.

    • Publisher confirms address publication from publisher to RabbitMQ.
    • Consumer acknowledgments address processing from RabbitMQ to consumer.

    RabbitMQ rule: Confirming publication does not prove that a consumer completed the business operation. Publisher confirms and consumer acknowledgments are separate mechanisms.

    RabbitMQ Prefetch

    Prefetch limits the number of unacknowledged deliveries sent to a consumer or channel, according to the client and broker configuration.

    Consumer prefetch:
    
    10
    
    
    Maximum assigned but
    unacknowledged deliveries:
    
    10

    A very large prefetch can overload a slow consumer and distribute work unevenly. A very small prefetch can reduce throughput for fast consumers. Tune it using measured processing duration and memory consumption.

    RabbitMQ Competing Consumers

    Task Queue
       |
       +-- Worker A
       +-- Worker B
       +-- Worker C

    Consumers connected to one queue compete for its messages. If several independent services need the same published message, the exchange can route copies to separate queues.

    RabbitMQ Strengths

    • Flexible exchange and binding model
    • Strong fit for task and worker queues
    • Competing-consumer processing
    • Publisher confirms and consumer acknowledgments
    • Suitable for direct, topic, fanout, and header-based routing
    • Useful when the broker should make routing decisions

    RabbitMQ Trade-offs

    • Traditional queues are not primarily designed as retained event history
    • Each independent subscriber commonly needs its own queue
    • Broker topology and queue lifecycle require active management
    • Incorrect acknowledgment or requeue behaviour can cause loss or repeated delivery
    • Poison messages can create repeated redelivery
    • Applications must tune prefetch and consumer concurrency

    Amazon SQS

    Amazon SQS is a managed AWS queue service. Producers send messages to a queue, and consumers poll the queue for available messages.

    Producer
        |
        v
    Amazon SQS Queue
        |
        +-- Consumer A
        +-- Consumer B
        +-- Consumer C

    Amazon SQS provides two queue types:

    • Standard queues
    • FIFO queues

    SQS Standard Queues

    Standard queues are designed for very high throughput, at-least-once delivery, and best-effort ordering.

    Applications using Standard queues must handle:

    • Occasional duplicate delivery
    • Messages arriving in a different order
    • Idempotent processing
    • Visibility timeout expiration
    • Explicit message deletion after success

    SQS FIFO Queues

    FIFO queues are designed for workloads where ordering and duplicate suppression within the queue's documented processing model are important.

    FIFO queues use message groups to provide separate ordered sequences inside one queue.

    FIFO Queue
    
    
    Message Group: learner-1042
    
    RegisterAccount
    EnrollInCourse
    CompleteCourse
    
    
    Message Group: learner-2055
    
    RegisterAccount
    EnrollInCourse
    CompleteCourse

    Messages inside one group are processed in order. Different message groups can support parallel progress.

    SQS Standard vs FIFO

    Area Standard Queue FIFO Queue
    Ordering Best effort FIFO semantics within the applicable message-group scope
    Delivery At least once; duplicates can occur Designed to prevent duplicate messages from being introduced into the queue under its deduplication model
    Parallelism Consumers process available messages broadly Parallelism comes from several independent message groups
    Primary fit High-volume work that tolerates duplicates and reordering Work requiring ordered processing within a business key

    SQS Visibility Timeout

    When a consumer receives an SQS message, the message becomes temporarily invisible to other consumers.

    Consumer receives message
          |
          v
    Visibility timeout begins
          |
          +-- Processing succeeds:
          |      consumer deletes message
          |
          +-- Message is not deleted:
                 visibility timeout expires
                 message becomes available again

    If the visibility timeout is shorter than normal processing duration, another consumer can receive the same message before the first consumer finishes.

    If the visibility timeout is unnecessarily long, failed work can take longer to become eligible for retry.

    SQS Delete after Success

    Receiving a message does not mean that processing has completed. The consumer should delete the message only after the required business outcome has completed successfully.

    Receive message
          |
          v
    Validate payload
          |
          v
    Process idempotently
          |
          v
    Commit business outcome
          |
          v
    Delete message

    SQS Dead-letter Queues

    An SQS queue can use a dead-letter queue for messages that are not processed successfully after the configured receive attempts.

    Source Queue
          |
          v
    Consumer processing fails
          |
          v
    Message becomes visible again
          |
          v
    Receive count reaches policy
          |
          v
    Move to Dead-letter Queue

    A dead-letter workflow should include monitoring, investigation, correction, and controlled redrive.

    SQS Strengths

    • Managed AWS queue infrastructure
    • Standard and FIFO queue options
    • Visibility timeout for in-progress work
    • Dead-letter queue support
    • Strong fit for AWS-native background processing
    • Consumers can scale independently from producers
    • No messaging-broker cluster administration by the application team

    SQS Trade-offs

    • It is a queue service rather than a replayable event-streaming log
    • Consumers poll queues rather than reading retained partition logs
    • Standard messages can be duplicated or delivered out of order
    • FIFO parallelism depends on appropriate message-group design
    • Complex fanout and routing can require additional AWS services or queues
    • Applications must tune visibility, retries, deletion, and dead-letter behaviour

    Detailed Feature Comparison

    Capability Kafka RabbitMQ SQS
    Event replay Core capability through retained records and offsets Not the primary model of traditional queues Deleted queue messages are not replayed normally
    Independent subscribers Separate consumer groups Separate routed queues Separate queues
    Task distribution Possible using one consumer group Core competing-consumer pattern Core queue pattern
    Broker-side routing Topics, keys, and partitions Rich exchanges and bindings Queue-level destination
    Processing completion Consumer commits offset according to processing design Consumer acknowledgment Consumer deletes message
    Temporary processing lock Partition assignment and consumer-group ownership Unacknowledged delivery state Visibility timeout
    Ordering Within a partition Depends on queue topology and processing model Best effort in Standard; FIFO behaviour in FIFO queues
    Failure isolation Topics, partitions, and consumer groups Virtual hosts, exchanges, queues, and consumers Separate queues and dead-letter queues
    Managed infrastructure Available through managed Kafka products Available through managed RabbitMQ products Native characteristic of the AWS service

    Choosing the Correct Platform

    Choose Kafka When

    • Events must remain replayable
    • Several independent consumer groups require the same event stream
    • Change Data Capture is a major requirement
    • Real-time stream processing is required
    • Ordering is needed within a stable business key
    • Historical events must rebuild projections or indexes
    • Event-stream retention is part of the architecture

    Choose RabbitMQ When

    • Messages require flexible exchange-based routing
    • Commands must be distributed among competing workers
    • Direct, topic, fanout, or header-based routing is useful
    • Per-message acknowledgment and requeue behaviour are important
    • The team can operate RabbitMQ or use an appropriate managed offering
    • Retained event replay is not the central requirement

    Choose Amazon SQS When

    • The solution runs primarily on AWS
    • A managed queue is preferred over operating brokers
    • The workload is background-task processing
    • Producer and consumer scaling should remain independent
    • Standard queue duplicates and reordering can be handled
    • FIFO message groups meet the required ordering scope
    • Visibility timeout and dead-letter queue semantics fit the workflow

    Selection rule: Select the platform from message meaning, replay requirements, routing complexity, ordering scope, consumer model, cloud environment, and operational responsibility. Do not select it only from benchmark numbers or product popularity.

    Using More than One Platform

    One organization can use different messaging technologies for different workloads.

    Business transaction
          |
          v
    Kafka event:
    
    CourseCompleted
          |
          +-- Analytics consumer group
          +-- Search consumer group
          +-- Certification service
                  |
                  v
            RabbitMQ or SQS task:
    
            GenerateCertificate
                  |
                  v
            Competing workers

    Kafka distributes and retains the business fact. RabbitMQ or SQS can then dispatch a concrete unit of work.

    Command and Event Examples

    Message Type Possible Platform Direction
    CourseCompleted Event Kafka when several independent consumers need the retained fact
    GenerateCertificate Command RabbitMQ or SQS when one worker should perform the task
    SendEnrollmentEmail Command RabbitMQ or SQS worker queue
    LearnerEnrolled Event Kafka when analytics, notifications, and search consume independently
    ResizeCourseImage Command RabbitMQ or SQS task queue

    Delivery and Idempotency

    All three systems require application-level attention to duplicate processing and uncertain outcomes.

    Consumer performs business update
          |
          v
    Update commits
          |
          v
    Acknowledgment, offset commit,
    or delete request fails
          |
          v
    Message can be processed again

    Use:

    • Stable message or event identifiers
    • Unique database constraints
    • Conditional writes
    • Processed-message records
    • Idempotency keys
    • Transactional consumer updates where supported

    Idempotent Message

    {
      "messageId": "stable-unique-message-id",
      "messageType": "GenerateCertificate",
      "messageVersion": 1,
      "tenantId": 17,
      "learnerId": 1042,
      "courseId": 42
    }

    Transactional Outbox

    A transactional outbox helps prevent a service from committing business data without recording the corresponding message.

    Local database transaction
          |
          +-- Update business record
          +-- Insert outbox message
          |
          v
    Commit once
          |
          v
    Publisher reads outbox
          |
          v
    Publish to Kafka,
    RabbitMQ, or SQS
          |
          v
    Mark publication progress

    The publisher can retry after failure. Consumers must still be idempotent because publication can occur more than once.

    Backpressure

    A messaging platform buffers work, but buffering does not create unlimited consumer capacity.

    Let \(P\) be the production rate and \(C\) be the successful consumer completion rate.

    Backlog grows when:

    \[ P > C \]

    A simplified backlog growth rate is:

    \[ BacklogGrowthRate = P - C \]

    Use:

    • Producer rate limits
    • Bounded consumer concurrency
    • Autoscaling based on backlog or lag
    • Batching where supported and appropriate
    • Load shedding for optional work
    • Dead-letter handling for poison messages

    Conceptual Kafka Configuration

    kafka:
      topic: learning-domain-events
      partitions: approved-partition-count
      replication: approved-replication-policy
    
      producer:
        key: trusted-aggregate-id
        acknowledgments: approved-durability-policy
    
      consumer:
        groupId: analytics-projection
        offsetCommit: after-durable-processing
    
      retention:
        policy: approved-replay-requirement
    
      observability:
        consumerLag: enabled
        partitionSkew: enabled
        rebalanceEvents: enabled
        retainedStorage: enabled

    Conceptual RabbitMQ Configuration

    rabbitmq:
      exchange:
        name: learning.commands
        type: topic
    
      queue:
        name: certificate-generation
        durable: true
    
      binding:
        routingKey: certificate.generate
    
      publisher:
        confirms: enabled
    
      consumer:
        acknowledgment: manual
        prefetch: approved-safe-limit
    
      failureHandling:
        boundedRetries: true
        deadLetterDestination: certificate-failures
    
      observability:
        readyMessages: enabled
        unacknowledgedMessages: enabled
        redeliveries: enabled

    Conceptual SQS Configuration

    sqs:
      queueType: standard-or-fifo
      queueName: certificate-generation
    
      processing:
        visibilityTimeout: approved-processing-boundary
        deleteAfterDurableSuccess: true
        idempotentConsumerRequired: true
    
      fifo:
        messageGroupKey: trusted-business-key
        deduplicationKey: stable-message-id
    
      failureHandling:
        deadLetterQueue: certificate-generation-failures
        maximumReceives: approved-limit
    
      observability:
        availableMessages: enabled
        inFlightMessages: enabled
        oldestMessageAge: enabled
        deadLetterMessages: enabled

    These are conceptual examples. Exact settings and supported guarantees must be verified against the deployed platform and version.

    PHP Idempotent Consumer

    <?php
    
    declare(strict_types=1);
    
    final class CertificateCommandHandler
    {
        public function __construct(
            private DatabaseConnection $database,
            private ProcessedMessageRepository $processedMessages,
            private CertificateService $certificates
        ) {
        }
    
        public function handle(
            array $message
        ): void {
            $messageId =
                (string)$message['messageId'];
    
            $this->database->transaction(
                function () use (
                    $messageId,
                    $message
                ): void {
                    if (
                        $this->processedMessages
                            ->exists(
                                consumerName:
                                    'certificate-command-handler',
    
                                messageId:
                                    $messageId
                            )
                    ) {
                        return;
                    }
    
                    $this->certificates
                        ->generateIfNotExists(
                            tenantId:
                                (int)$message['tenantId'],
    
                            learnerId:
                                (int)$message['learnerId'],
    
                            courseId:
                                (int)$message['courseId']
                        );
    
                    $this->processedMessages
                        ->record(
                            consumerName:
                                'certificate-command-handler',
    
                            messageId:
                                $messageId
                        );
                }
            );
        }
    }

    Broker acknowledgment, offset commit, or SQS deletion should happen only after this durable processing boundary succeeds.

    Security Considerations

    • Grant producers access only to required topics, exchanges, or queues.
    • Grant consumers access only to their required destinations.
    • Encrypt messaging traffic and stored messages according to policy.
    • Do not publish passwords, access tokens, or reusable credentials.
    • Use trusted tenant context rather than client-supplied routing identity.
    • Protect broker-management and queue-management operations.
    • Define event retention according to privacy and business requirements.
    • Redact sensitive payloads from processing logs and dead-letter diagnostics.
    • Audit administrative changes and message-access permissions.

    Observability

    Kafka Metrics

    • Records produced by topic and partition
    • Consumer-group lag
    • Oldest unprocessed-event age
    • Partition traffic skew
    • Consumer processing rate
    • Offset-commit failures
    • Consumer-group rebalances
    • Broker storage and replication health

    RabbitMQ Metrics

    • Ready-message count
    • Unacknowledged-message count
    • Publish and delivery rates
    • Consumer acknowledgment rate
    • Redelivery rate
    • Queue depth and message age
    • Dead-letter volume
    • Consumer and connection health

    SQS Metrics

    • Available-message count
    • In-flight-message count
    • Age of oldest message
    • Messages sent, received, and deleted
    • Receive count and retry behaviour
    • Dead-letter queue depth
    • Consumer completion rate
    • Visibility-timeout expirations inferred from redelivery

    Alert Conditions

    Alert when:

    • Kafka consumer lag continues increasing
    • One Kafka partition receives disproportionate traffic
    • Kafka consumers rebalance repeatedly
    • RabbitMQ ready or unacknowledged messages continue growing
    • RabbitMQ redelivery volume increases unexpectedly
    • SQS oldest-message age exceeds the processing objective
    • SQS messages repeatedly return after visibility timeout
    • A dead-letter queue receives unexpected messages
    • Consumer completion rate remains below producer rate
    • Poison messages block useful processing
    • Messaging storage or broker resources approach capacity
    • Schema compatibility errors increase

    Common Selection Mistakes

    1

    Choosing Kafka for Every Background Task

    A simple worker queue can be easier when replay and independent event consumers are not required.

    2

    Choosing RabbitMQ for Long-term Event Replay

    Traditional acknowledged queues are completion-oriented rather than retained-history oriented.

    3

    Using SQS Standard while Assuming Strict Order

    Standard queues use best-effort ordering and can deliver a message more than once.

    4

    Using One SQS FIFO Message Group for Every Message

    One group serializes ordered processing and can limit available parallelism.

    5

    Assuming Kafka Provides Global Topic Order

    Kafka ordering is scoped to each partition.

    6

    Adding Kafka Consumers beyond Partition Count

    Additional consumers in the same group can remain without an active partition assignment.

    7

    Confusing RabbitMQ Publisher Confirms with Consumer Completion

    Publisher confirms do not prove that a consumer completed the business operation.

    8

    Using Automatic Acknowledgment before RabbitMQ Processing

    A consumer crash can lose work that RabbitMQ already considers delivered.

    9

    Deleting an SQS Message before Business Commit

    The task can disappear before its intended outcome is durable.

    10

    Ignoring Idempotency

    Redelivery, replay, acknowledgment loss, or offset recovery can repeat a business operation.

    11

    Using Unbounded Retries

    Poison messages consume capacity and prevent useful work from progressing.

    12

    Selecting from Throughput Claims Alone

    Replay, routing, ordering, management, cost, and failure recovery can be more important than raw message rate.

    Recommended Test Cases

    Test Expected Evidence
    Kafka consumer restart The consumer resumes according to its committed offset
    Kafka replay A projection can be rebuilt without duplicate external effects
    Kafka hot partition Partition skew is detected and protected
    Kafka group rebalance Partition ownership changes without silently skipping required records
    RabbitMQ consumer failure An unacknowledged message is requeued or redelivered according to policy
    RabbitMQ lost acknowledgment Idempotency prevents a duplicate business effect
    RabbitMQ routing Direct, topic, or fanout rules deliver messages only to intended queues
    RabbitMQ poison message Bounded attempts move the message to the approved failure path
    SQS visibility timeout A failed consumer causes the message to become available again
    SQS duplicate delivery Consumer idempotency prevents a duplicate outcome
    SQS Standard ordering The application remains correct when messages arrive out of order
    SQS FIFO groups Messages remain ordered within each group while separate groups progress independently
    Dead-letter workflow Failed messages are observable, diagnosable, and safely redriven
    Producer surge Backlog controls preserve broker, consumer, and downstream stability
    Schema evolution Old and new consumers remain compatible during deployment

    Best Practices

    Recommended Practices

    • Choose the platform from the messaging model, not popularity.
    • Use Kafka for retained, replayable event streams.
    • Use RabbitMQ for flexible routing and competing task consumers.
    • Use SQS for managed AWS-native queues.
    • Separate commands from events.
    • Give every message a stable unique identifier.
    • Make all consumers idempotent.
    • Commit offsets, acknowledge deliveries, or delete messages only after durable processing.
    • Use a stable Kafka partition key for required per-entity ordering.
    • Avoid Kafka hot partitions.
    • Tune RabbitMQ prefetch from measured consumer behaviour.
    • Use RabbitMQ publisher confirms and consumer acknowledgments for their separate purposes.
    • Set SQS visibility according to processing requirements.
    • Use several SQS FIFO message groups when ordered streams can progress independently.
    • Use bounded retries with exponential backoff and jitter.
    • Use dead-letter handling for poison messages.
    • Use the transactional outbox for reliable publication.
    • Version message and event schemas.
    • Monitor backlog age and consumer lag, not count alone.
    • Test replay, redelivery, rebalancing, visibility expiry, and recovery.

    Practice Exercise

    Select Kafka, RabbitMQ, or Amazon SQS for the asynchronous workflows in your online learning platform.

    Requirements

    1. Publish learner-domain events to Kafka.
    2. Create independent event consumers for analytics and search.
    3. Use learner ID as the Kafka key where per-learner ordering is required.
    4. Replay retained events to rebuild a search index.
    5. Create a RabbitMQ topic exchange for complex notification routing.
    6. Create separate RabbitMQ queues for email and push notifications.
    7. Configure manual acknowledgments and bounded prefetch.
    8. Use SQS for AWS-native course-report generation.
    9. Compare SQS Standard and FIFO for the report workflow.
    10. Create an SQS dead-letter queue.
    11. Add idempotency to every consumer.
    12. Use a transactional outbox for event and command publication.
    13. Test poison-message handling on all selected platforms.
    14. Monitor Kafka lag, RabbitMQ queue depth, and SQS oldest-message age.
    15. Document why each workflow uses its selected platform.

    Selection Template

    Workflow Suggested Platform Direction Reason
    Domain-event distribution Kafka Independent consumers and replayable history
    Search-index update stream Kafka Consumer can replay after index recreation
    Complex topic-based notification routing RabbitMQ Exchange and binding model supports flexible routing
    Competing media-processing workers RabbitMQ or SQS Each task requires one successful worker
    AWS-native serverless background job SQS Managed AWS queue with worker decoupling
    Ordered AWS task stream per learner SQS FIFO Message groups can preserve per-learner ordering

    Frequently Asked Questions

    1

    Is Kafka a traditional message queue?

    Kafka is primarily a distributed event-streaming log. It can distribute work through consumer groups, but records remain according to retention rather than being removed when one consumer reads them.

    2

    What is RabbitMQ best suited for?

    RabbitMQ is well suited to traditional task queues and routing messages through exchanges and bindings to one or more queues.

    3

    What is Amazon SQS best suited for?

    Amazon SQS is well suited to AWS-native applications requiring a managed queue for asynchronous task processing and service decoupling.

    4

    Which platform supports replay most naturally?

    Kafka supports replay through retained topic records and consumer offsets.

    5

    Which platform has the richest broker-side routing?

    RabbitMQ provides exchanges and bindings supporting direct, topic, fanout, header-based, and additional routing behaviours.

    6

    What is an SQS visibility timeout?

    It is the period during which a received message remains hidden from other consumers while one consumer processes it.

    7

    Does Kafka guarantee global ordering?

    No. Ordering is provided within each partition rather than across an entire multi-partition topic.

    8

    Does SQS Standard guarantee ordering?

    No. Standard queues provide best-effort ordering and applications must tolerate messages arriving out of order.

    9

    How does RabbitMQ know that processing succeeded?

    A consumer sends an acknowledgment after reaching its successful processing boundary.

    10

    Why do consumers still need idempotency?

    A failure after the business operation but before offset commit, acknowledgment, or message deletion can cause processing to occur again.

    11

    Can these platforms be used together?

    Yes. Kafka can distribute retained domain events, while RabbitMQ or SQS dispatches concrete tasks derived from those events.

    12

    Which one should I choose?

    Choose Kafka for retained event streaming and replay, RabbitMQ for flexible broker routing and task queues, or SQS for managed AWS-native queuing. Validate the final choice against ordering, delivery, throughput, security, cost, and operational requirements.

    Key Takeaway

    Kafka, RabbitMQ, and Amazon SQS all support asynchronous communication, but they use different models. Kafka stores retained records in partitioned topics and lets independent consumer groups track offsets and replay events. RabbitMQ routes producer messages through exchanges and bindings into queues, where consumers acknowledge completed deliveries. Amazon SQS provides managed Standard and FIFO queues, using visibility timeout and explicit deletion to control processing. Choose Kafka for event streams, Change Data Capture, replay, analytics, and multiple independent subscribers. Choose RabbitMQ for flexible routing and competing worker queues. Choose SQS for AWS-native task processing without operating broker infrastructure. Regardless of platform, use idempotent consumers, transactional publication, bounded retries, dead-letter handling, schema versioning, backpressure, least-privilege access, and workload-specific observability.