Table of Contents

    batching

    MESSAGING & ASYNCHRONOUS PROCESSING

    Batching

    Learn how batching groups multiple messages or operations into one processing unit to improve throughput and reduce network, storage, and transaction overhead. Understand batch size, flush intervals, latency trade-offs, partial failures, ordering, retries, idempotency, Kafka consumer batches, RabbitMQ acknowledgments, SQS batch APIs, database bulk writes, adaptive batching, backpressure, and observability.

    Introduction

    A messaging consumer can process one message at a time or process several messages together as a batch.

    One-message processing:
    
    Message 1 -> Network call -> Database commit
    
    Message 2 -> Network call -> Database commit
    
    Message 3 -> Network call -> Database commit

    Processing each message separately is simple, but every message can create its own network request, database transaction, acknowledgment, serialization operation, and application overhead.

    Batch processing:
    
    Message 1
    Message 2
    Message 3
    Message 4
    Message 5
        |
        v
    One batch operation
        |
        v
    One bulk database write
    or one downstream request

    Batching improves efficiency by sharing fixed processing costs across several records. However, a consumer must wait until enough records are available or until a flush timer expires.

    Therefore, batching creates an important system-design trade-off:

    • Larger batches can improve throughput.
    • Smaller batches can reduce waiting latency.
    • Oversized batches can increase memory use, transaction duration, retry cost, and failure impact.

    Core idea: Batching amortizes fixed processing overhead across several records. A good batch policy improves throughput without violating latency, memory, ordering, retry, transaction, or downstream capacity requirements.

    Prerequisites

    # Prerequisite Why It Is Needed
    1 Queues vs logs Queue and log consumers expose different batch-fetching and progress models.
    2 Kafka, RabbitMQ, and Amazon SQS Each platform provides different batching controls.
    3 Consumer groups Kafka consumers receive batches from their assigned partitions.
    4 Delivery semantics A failed batch can cause some or all records to be delivered again.
    5 Retries and DLQs Partial batch failures need a safe retry and terminal-failure policy.
    6 Ordering Parallel or partial batch processing can change completion order.
    7 Backpressure Batching must remain within consumer and downstream capacity.

    What Is Batching?

    Batching is the practice of collecting several messages, records, requests, or updates and processing them together as one logical unit.

    Incoming records
          |
          v
    Batch accumulator
          |
          +-- Record count reached
          |
          +-- Byte limit reached
          |
          +-- Flush timer expired
          |
          v
    Process accumulated batch

    A batch can be created at several layers:

    • Producer batching
    • Broker or client batching
    • Consumer fetching
    • Application processing
    • Database bulk writing
    • External API submission
    • Acknowledgment or offset commitment
    Batch Lifecycle
    collect records → reach size or time boundary → validate batch → process records → handle partial failures → record progress

    Individual vs Batch Processing

    Area Individual Processing Batch Processing
    Fixed overhead Paid for every record Shared by several records
    Latency Processing can begin immediately Records can wait for the batch to fill or flush
    Throughput Can be limited by repeated calls and commits Can improve through bulk operations
    Memory Usually lower per processing unit Increases as records accumulate
    Failure scope Normally limited to one record One failure can affect several records
    Retry cost Retry one record Can require retrying or splitting a batch
    Ordering Simpler to reason about Requires control when records are processed in parallel

    Why Batching Improves Throughput

    Many operations contain a fixed cost plus a per-record cost.

    A simplified model is:

    \[ BatchCost = FixedCost + \left( RecordCount \times PerRecordCost \right) \]

    The average cost per record is:

    \[ AverageCostPerRecord = \frac{ FixedCost }{ RecordCount } + PerRecordCost \]

    As the number of records in the batch grows, the fixed cost is divided among more records.

    Fixed operation cost:
    
    One network round trip
    
    
    Per-record cost:
    
    Serialization and validation
    
    
    Batching result:
    
    One network round trip carries
    several serialized records.

    Throughput vs Latency

    A batch normally waits until it reaches a size boundary or a time boundary.

    Large batch:
    
    Higher throughput potential
    Higher waiting latency
    Higher memory use
    
    
    Small batch:
    
    Lower waiting latency
    Lower failure scope
    More fixed overhead

    The correct batch size depends on the business latency objective and the capacity of the consumer and downstream systems.

    Trade-off: Do not optimize batch size for maximum throughput alone. Include waiting time, processing time, retry time, memory, transaction duration, downstream limits, and recovery behaviour.

    Batch Boundaries

    Boundary Meaning
    Record count Flush when the batch contains a configured number of records
    Byte count Flush when the total payload size reaches a configured limit
    Time interval Flush when the oldest waiting record reaches the configured interval
    Business grouping Flush when a complete logical group is available
    Partition boundary Keep records separated according to their source partitions
    Transaction boundary Limit the records included in one atomic database transaction
    Resource-pressure boundary Flush or shrink the batch when memory or downstream capacity is constrained

    Count-Based Batching

    Maximum batch count:
    
    100 records
    
    
    Consumer accumulates:
    
    Record 1
    Record 2
    ...
    Record 100
    
    
    Record count reached
          |
          v
    Process batch

    Count-based batching is easy to understand, but records can have different payload sizes and processing costs.

    Size-Based Batching

    Maximum batch size:
    
    Configured byte limit
    
    
    Record A:
    
    Small payload
    
    
    Record B:
    
    Large payload
    
    
    Flush when accumulated bytes
    reach the safe boundary.

    Byte-based limits protect messaging clients, network calls, memory, storage, and downstream APIs from oversized requests.

    Time-Based Batching

    Flush interval:
    
    Configured time boundary
    
    
    Low traffic:
    
    Only 5 records arrive
    
    
    Timer expires
          |
          v
    Process 5-record batch
    rather than waiting indefinitely.

    A flush timer prevents low-volume messages from waiting forever for a count threshold that might not be reached.

    Hybrid Batching

    A common policy flushes when any approved boundary is reached.

    Flush when:
    
    - Maximum record count is reached
    
    OR
    
    - Maximum byte size is reached
    
    OR
    
    - Maximum waiting time is reached

    Hybrid batching balances throughput, payload size, memory, and latency.

    Producer Batching

    Producer batching groups records before sending them to the messaging system.

    Application records
          |
          v
    Producer accumulator
          |
          v
    One broker request
    containing several records

    Producer batching can reduce:

    • Network round trips
    • Protocol overhead
    • Serialization setup cost
    • Broker request count

    Producer batching can increase the time that the first record waits before transmission.

    Consumer Batching

    Consumer batching retrieves or accumulates several messages before invoking business processing.

    Messaging destination
          |
          v
    Consumer fetches several records
          |
          v
    Validate records
          |
          v
    Process as a batch
          |
          v
    Record individual
    or batch progress

    Consumer batch size must remain within processing-time, memory, poll, visibility, acknowledgment, and downstream transaction limits.

    Kafka Batching

    Kafka producers can send records efficiently in grouped requests, and Kafka consumers can retrieve several records during a fetch and poll cycle.

    Kafka Topic Partitions
          |
          v
    Consumer poll
          |
          v
    ConsumerRecords batch
          |
          v
    Application processing
          |
          v
    Offset progress

    Internal architecture material also describes AWS Lambda reading Kafka records in batches and passing those batches to function code for processing. 【1-fb6756】【2-ea8918】

    Kafka Batch Concerns

    • Records can come from several assigned partitions.
    • Offsets must be committed only for records that reached the intended durable boundary.
    • One failing record can complicate progress for later records in the same partition.
    • Large batches can exceed safe consumer-processing time.
    • Parallel batch processing can alter per-partition completion order.
    • One hot partition can dominate the batch.

    Kafka Offset Handling

    Fetched offsets:
    
    100
    101
    102
    103
    104
    
    
    Processed successfully:
    
    100
    101
    
    
    Offset 102 fails
    
    
    Offsets 103 and 104:
    
    Require an ordering-aware
    processing decision.

    Committing beyond a failed record can skip required processing after restart. Reprocessing the complete batch can repeat already completed records.

    Use idempotency and partition-aware progress tracking.

    RabbitMQ Batching

    RabbitMQ consumers can receive several deliveries over time and process them together in the application.

    RabbitMQ Queue
          |
          v
    Consumer receives deliveries
          |
          v
    Application forms bounded batch
          |
          v
    Bulk process
          |
          v
    Acknowledge successful deliveries

    Consumer prefetch affects how many unacknowledged deliveries can be assigned under the applicable configuration. Internal learning material identifies durable queues and prefetch as important RabbitMQ processing controls. 【3-3222f4】

    RabbitMQ Batch Concerns

    • Do not acknowledge deliveries before durable processing succeeds.
    • Large prefetch values can accumulate excessive in-flight work.
    • A partial batch failure requires per-message acknowledgment or retry decisions.
    • Multiple consumers can change completion order.
    • Immediate requeueing can create rapid retry loops.

    Amazon SQS Batching

    Amazon SQS clients can send, receive, and delete more than one message using supported batch-oriented API operations.

    SQS Queue
          |
          v
    Consumer receives message batch
          |
          v
    Process each batch entry
          |
          +-- Successful entries:
          |      delete safely
          |
          +-- Failed entries:
                 retry according
                 to message policy

    SQS Batch Concerns

    • A batch request can contain successful and failed entries.
    • Retry only entries that did not complete successfully.
    • Each message still requires idempotent business processing.
    • Visibility timeout must cover the intended processing boundary.
    • Large consumer batches can increase duplicate-processing risk when visibility expires.
    • FIFO message groups require ordering-aware batch handling.

    SQS batch rule: Treat each message entry as an independent business operation unless the business workflow explicitly defines one atomic batch.

    Database Batching

    Consumers often batch records to reduce database round trips.

    Without bulk operation:
    
    100 records
        -> 100 database calls
    
    
    With bulk operation:
    
    100 records
        -> bounded bulk statement
        or controlled transaction

    Possible database techniques include:

    • Multi-row insert
    • Bulk-copy operation
    • Set-based update
    • Upsert using a staging table
    • Stored procedure accepting structured input
    • Bounded transaction chunks

    Multi-Row Insert

    INSERT INTO progress_events
    (
        event_id,
        tenant_id,
        learner_id,
        course_id,
        event_type,
        source_version
    )
    VALUES
    (
        :event_id_1,
        :tenant_id_1,
        :learner_id_1,
        :course_id_1,
        :event_type_1,
        :source_version_1
    ),
    (
        :event_id_2,
        :tenant_id_2,
        :learner_id_2,
        :course_id_2,
        :event_type_2,
        :source_version_2
    );

    Exact limits and optimal bulk-write methods depend on the selected database, driver, schema, indexes, transaction log, and statement-size constraints.

    Transaction Size

    One large transaction can hold locks, consume transaction-log capacity, and create expensive rollback work.

    Very large batch
          |
          v
    Long transaction
          |
          +-- Longer lock duration
          +-- More transaction-log use
          +-- Larger rollback cost
          +-- Increased contention
          +-- Longer failure recovery

    Divide very large workloads into bounded chunks unless the complete operation must be atomic.

    Atomic vs Non-Atomic Batch

    Model Behaviour Main Concern
    Atomic batch Every record succeeds or the complete transaction rolls back One bad record can fail the complete batch
    Non-atomic batch Records can succeed or fail independently Requires per-record result tracking
    Chunked atomic processing Each bounded chunk is transactional The complete workload can be partially completed
    Best-effort batch Process as much as possible and report failures Requires reconciliation and clear correctness rules

    Partial Batch Failure

    Batch contains:
    
    Record A -> success
    
    Record B -> success
    
    Record C -> failure
    
    Record D -> success
    
    Record E -> failure

    The consumer must decide whether to:

    • Roll back the complete batch
    • Commit successful records and retry failed records
    • Split the batch and retry smaller groups
    • Send permanently invalid records to a DLQ
    • Stop later records when ordering requires it

    Batch Splitting

    Batch splitting can isolate the record causing a repeatable failure.

    Batch of 8 fails
          |
          v
    Split into two batches of 4
          |
          +-- First half succeeds
          |
          +-- Second half fails
                  |
                  v
            Split failed half again

    This technique is useful when the downstream system reports only that the batch failed and does not identify the failing record.

    Batch splitting should remain bounded. A deterministic invalid record should eventually reach the terminal failure path rather than causing unlimited subdivision and retry.

    Batch Retries

    Retrying a complete batch can repeat records that already succeeded.

    Batch processing:
    
    A -> success
    B -> success
    C -> uncertain failure
    
    
    Retry complete batch:
    
    A -> repeated
    B -> repeated
    C -> retried

    Use stable per-record identifiers and idempotent processing even when the transport operation is batched.

    Per-Record Idempotency

    {
      "batchId": "stable-batch-id",
      "records": [
        {
          "messageId": "stable-message-id-1",
          "operation": "UpdateProgress"
        },
        {
          "messageId": "stable-message-id-2",
          "operation": "UpdateProgress"
        }
      ]
    }

    The batch ID identifies the transport group. Each record ID identifies its business operation.

    Idempotency rule: Do not rely only on the batch ID for duplicate protection. Preserve a stable identity for every independently meaningful operation inside the batch.

    Batching and Ordering

    A batch can contain records from several business entities or partitions.

    Batch:
    
    Learner A, Sequence 10
    
    Learner B, Sequence 4
    
    Learner A, Sequence 11
    
    Learner C, Sequence 8

    Independent entities can be processed concurrently. Records belonging to the same ordered entity must preserve the required sequence.

    Possible controls include:

    • Group records by business key
    • Process each key sequentially
    • Process different keys concurrently
    • Use sequence numbers or source versions
    • Prevent an older retry from overwriting newer state

    Per-Partition Batch Processing

    Kafka poll returns:
    
    Partition 0:
    Offsets 100, 101, 102
    
    Partition 1:
    Offsets 50, 51, 52
    
    
    Processing policy:
    
    Preserve order inside each partition
    
    Allow Partition 0 and Partition 1
    to progress independently.

    This preserves partition-local ordering while retaining useful parallelism.

    Memory Management

    A batch occupies memory until it is processed or discarded.

    A simplified estimate is:

    \[ ApproximateBatchMemory = RecordCount \times AverageRecordSize + ProcessingOverhead \]

    Actual memory includes deserialized objects, indexes, temporary structures, library buffers, output records, and garbage-collection overhead.

    Batching and Backpressure

    Increasing batch size can improve throughput until memory, processing time, transaction duration, or downstream capacity becomes the bottleneck.

    Pressure increases
          |
          v
    Reduce batch size
    or consumer concurrency
          |
          v
    Protect downstream capacity
    
    
    Capacity is healthy
          |
          v
    Increase batch size gradually
    within approved limits

    Batch size and concurrent batch count must be considered together.

    A simplified in-flight estimate is:

    \[ InFlightRecords = BatchSize \times ConcurrentBatches \]

    Adaptive Batching

    Adaptive batching changes batch size or flush behaviour according to observed load and processing capacity.

    Low traffic:
    
    Use small batches
    Flush quickly
    
    
    High traffic and healthy capacity:
    
    Use larger batches
    Improve throughput
    
    
    Downstream saturation:
    
    Reduce concurrency
    Shrink batches
    Slow intake

    Adaptive logic should use bounded and tested limits. Continuous adjustment without stabilization can cause oscillation.

    Conceptual Batch Policy

    batching:
      enabled: true
    
      boundaries:
        maximumRecords: approved-record-limit
        maximumBytes: approved-byte-limit
        maximumWait: approved-latency-boundary
    
      processing:
        atomicity: workload-specific
        concurrency: approved-safe-limit
        partitionAware: true
        orderingAware: true
    
      failureHandling:
        perRecordResults: required
        boundedRetries: true
        splitFailedBatch: workload-specific
        deadLetterDestination: approved-dlq
    
      idempotency:
        batchId: required
        perRecordMessageId: required
    
      backpressure:
        reduceBatchOnSaturation: true
        boundedInFlightRecords: true
    
      observability:
        actualBatchSize: enabled
        batchFillTime: enabled
        batchDuration: enabled
        partialFailures: enabled
        retriedRecords: enabled

    This is a conceptual policy. Exact limits must be measured against the messaging client, broker, database, network, runtime, and business latency requirements.

    Java Batch Accumulator

    public final class BatchAccumulator<T> {
    
        private final int maximumRecords;
        private final Duration maximumWait;
        private final List<T> records;
        private Instant batchStartedAt;
    
        public BatchAccumulator(
            int maximumRecords,
            Duration maximumWait
        ) {
            this.maximumRecords =
                maximumRecords;
    
            this.maximumWait =
                maximumWait;
    
            this.records =
                new ArrayList<>();
        }
    
        public synchronized Optional<List<T>> add(
            T record
        ) {
            if (records.isEmpty()) {
                batchStartedAt =
                    Instant.now();
            }
    
            records.add(
                record
            );
    
            if (
                records.size()
                    >= maximumRecords
            ) {
                return Optional.of(
                    drain()
                );
            }
    
            return Optional.empty();
        }
    
        public synchronized Optional<List<T>>
            flushIfExpired() {
    
            if (
                records.isEmpty()
                || batchStartedAt == null
            ) {
                return Optional.empty();
            }
    
            Duration age =
                Duration.between(
                    batchStartedAt,
                    Instant.now()
                );
    
            if (
                age.compareTo(maximumWait)
                    >= 0
            ) {
                return Optional.of(
                    drain()
                );
            }
    
            return Optional.empty();
        }
    
        private List<T> drain() {
            List<T> batch =
                List.copyOf(
                    records
                );
    
            records.clear();
            batchStartedAt = null;
    
            return batch;
        }
    }

    This educational example does not include byte limits, graceful shutdown, timer scheduling, concurrency safety outside the accumulator, message acknowledgments, retries, metrics, or persistence of records waiting in memory.

    PHP Batch Processor

    <?php
    
    declare(strict_types=1);
    
    final class BatchProcessor
    {
        public function __construct(
            private RecordValidator $validator,
            private BulkRepository $repository,
            private FailurePublisher $failurePublisher
        ) {
        }
    
        /**
         * @param list<array<string, mixed>> $records
         */
        public function process(
            array $records
        ): void {
            $validRecords = [];
    
            foreach ($records as $record) {
                $validationResult =
                    $this->validator->validate(
                        $record
                    );
    
                if (
                    !$validationResult->isValid()
                ) {
                    $this->failurePublisher
                        ->publish(
                            $record,
                            $validationResult
                        );
    
                    continue;
                }
    
                $validRecords[] =
                    $record;
            }
    
            if ($validRecords === []) {
                return;
            }
    
            $this->repository
                ->upsertIdempotently(
                    $validRecords
                );
        }
    }

    Production processing requires transactional rules, per-record results, bounded retries, acknowledgment logic, ordering controls, schema validation, authorization, metrics, and safe handling when the bulk repository partially succeeds.

    Learning-platform Examples

    Workflow Batching Direction Main Concern
    Learner-progress events Bulk upsert a bounded group of progress changes Preserve per-learner version order
    Certificate generation Fetch several eligible tasks but generate each certificate idempotently One invalid task must not duplicate successful certificates
    Email delivery Submit a provider-supported bounded batch Track delivery result for every notification
    Search indexing Bulk index course and article changes Retry only failed documents and reject stale versions
    Analytics ingestion Write event batches to analytical storage Balance throughput with event freshness
    Report generation Process records in bounded chunks Avoid large transactions and excessive memory

    Security and Tenant Isolation

    • Do not combine incompatible tenant data in an unauthorized batch.
    • Derive tenant identity from trusted message context.
    • Apply authorization before bulk processing.
    • Do not let one tenant create an unbounded batch.
    • Protect batch payloads and failure metadata.
    • Use least-privilege messaging and database identities.
    • Audit bulk redrive and replay operations.
    • Do not log complete sensitive batch payloads.

    Observability

    Useful batch-processing metrics include:

    • Configured maximum batch size
    • Actual records per batch
    • Actual bytes per batch
    • Batch fill time
    • Batch processing duration
    • Records processed per second
    • Successful records per batch
    • Failed records per batch
    • Partial batch failures
    • Retried records
    • Dead-lettered records
    • Memory used by batches
    • Database transaction duration
    • Consumer lag or queue age
    • Downstream throttling

    Structured Batch Event

    {
      "batchId": "stable-batch-id",
      "consumer": "progress-projection",
      "recordCount": "measured-count",
      "successfulRecords": "measured-count",
      "failedRecords": "measured-count",
      "batchResult": "partially-completed",
      "processingDuration": "measured-duration"
    }

    Avoid recording complete batch payloads when identifiers, counts, safe error codes, and protected references are sufficient.

    Alert Conditions

    Alert when:

    • Batch processing duration continues increasing
    • Actual batches remain unusually small during high traffic
    • Records wait too long for batch formation
    • Partial batch failures increase
    • Complete batch retries increase
    • Memory pressure grows with batch size
    • Database transaction duration exceeds the safe objective
    • Consumer lag grows despite large batches
    • Visibility or polling boundaries are exceeded
    • Failed records repeatedly return to the same batch path
    • Batch processing changes required message order
    • Downstream bulk APIs throttle or reject batches

    Troubleshooting Workflow

    1. Identify the producer, queue or topic, and consumer.
    2. Check the configured record, byte, and time boundaries.
    3. Measure actual records and bytes in each batch.
    4. Measure batch fill time and processing time separately.
    5. Check memory and in-flight record count.
    6. Check downstream transaction or API limits.
    7. Check whether failures are complete-batch or per-record failures.
    8. Check retry and DLQ behaviour.
    9. Check acknowledgment, deletion, or offset timing.
    10. Check per-partition and per-key ordering.
    11. Check duplicate and idempotency records.
    12. Reduce batch size or concurrency when downstream saturation increases.
    13. Increase batch size gradually when fixed overhead is the confirmed bottleneck.

    Common Batching Mistakes

    1

    Using Only a Count Boundary

    A batch containing large records can exceed memory, network, or API limits.

    2

    Waiting Forever for a Full Batch

    Low-volume messages miss their latency objective because no flush timer exists.

    3

    Making the Batch Too Large

    Memory use, transaction duration, rollback cost, and retry impact increase.

    4

    Acknowledging before the Batch Is Durable

    Consumer failure can permanently lose records that were not committed.

    5

    Retrying the Complete Batch Blindly

    Successful records are processed again along with the failed records.

    6

    Using Only a Batch-Level Idempotency Key

    Independently meaningful records lack duplicate protection after a partial retry.

    7

    Ignoring Partial Success

    The consumer treats the complete request as failed even though some records already changed downstream state.

    8

    Parallelizing Ordered Records

    Records are received in order but complete in another sequence.

    9

    Using One Large Database Transaction

    Lock duration, contention, logging, and rollback work become excessive.

    10

    Increasing Batch Size during Downstream Saturation

    Larger requests add pressure to a dependency that is already overloaded.

    11

    Ignoring Shutdown Flush

    Records waiting in an in-memory partial batch can be lost during shutdown.

    12

    Measuring Throughput but Not Waiting Latency

    Total throughput improves while individual messages wait too long for batch formation.

    Recommended Test Cases

    Test Expected Evidence
    Maximum record count The batch flushes at the configured record boundary.
    Maximum byte size Large records cause an earlier safe flush.
    Low traffic The flush timer processes a partially filled batch.
    Complete batch success All records reach their durable processing boundary.
    One invalid record The approved partial-failure policy is applied.
    Uncertain downstream response Per-record idempotency prevents duplicate business effects.
    Database transaction failure The complete atomic chunk rolls back or is retried safely.
    Kafka mixed partitions Offset progress remains correct for every partition.
    RabbitMQ batch failure Only successfully processed deliveries are acknowledged according to policy.
    SQS partial result Only failed entries remain eligible for retry.
    Ordered entity records Records for the same key commit in the required sequence.
    Graceful shutdown A partial batch is flushed or safely preserved.
    Downstream saturation Batch size or concurrency decreases within bounded policy.
    Poison record Bounded splitting or retry isolates the record in the DLQ.

    Batching Best Practices

    Recommended Practices

    • Use batching when fixed per-call overhead is a measured bottleneck.
    • Define record-count, byte-size, and maximum-wait boundaries.
    • Use a flush timer for low-volume traffic.
    • Keep batches within memory and downstream payload limits.
    • Measure batch fill time and processing time separately.
    • Use bounded transaction chunks.
    • Define whether the batch is atomic or partially successful.
    • Track results for every independently meaningful record.
    • Assign a stable message ID to every record.
    • Make every record idempotent.
    • Retry only failed or uncertain records where possible.
    • Use bounded batch splitting when the failed record is unknown.
    • Send permanent failures to a DLQ.
    • Preserve ordering per partition or business key.
    • Commit offsets, acknowledgments, or deletions only after durable processing.
    • Limit the number of concurrent batches.
    • Reduce batch size or concurrency during downstream saturation.
    • Flush or preserve partial batches during graceful shutdown.
    • Protect tenant isolation inside every batch.
    • Test complete failure, partial failure, retry, shutdown, and recovery.

    Practice Exercise

    Design batch processing for learner-progress events in your online learning platform.

    Requirements

    1. Read progress events from a partitioned Kafka topic.
    2. Use learner ID and course ID as the ordering key.
    3. Define maximum record count, byte size, and wait time.
    4. Group records by source partition and business key.
    5. Bulk upsert the validated progress records.
    6. Add stable event IDs and source versions.
    7. Prevent old events from overwriting newer progress.
    8. Track successful and failed records separately.
    9. Retry only temporary failures.
    10. Move invalid records to a DLQ.
    11. Commit Kafka progress only after durable database processing.
    12. Simulate one invalid record in a batch.
    13. Simulate an uncertain database response.
    14. Test graceful shutdown with a partially filled batch.
    15. Compare throughput and latency using several bounded batch sizes.

    Batching-design Template

    Decision Selected Direction Risk Controlled
    Maximum count Measured bounded record count Limits batch processing and retry scope
    Maximum bytes Measured payload-size boundary Prevents oversized memory and network requests
    Maximum wait Business latency objective Prevents low-volume messages from waiting indefinitely
    Atomicity Chunk-level database transaction Limits rollback and lock duration
    Partial failure Per-record result and retry classification Prevents repeated processing of successful records
    Idempotency Stable ID for each record Prevents duplicate business effects
    Ordering Sequential per key, parallel across keys Preserves correctness without removing all parallelism
    Backpressure Bounded concurrent batches Protects downstream databases and services

    Frequently Asked Questions

    1

    What is batching?

    Batching collects several messages or operations and processes them together as one bounded processing unit.

    2

    Why does batching improve throughput?

    It shares fixed network, protocol, transaction, and application overhead across several records.

    3

    Does batching always improve performance?

    No. Oversized batches can increase latency, memory use, transaction duration, retry cost, and failure impact.

    4

    How should a batch be flushed?

    A common design flushes when a record-count, byte-size, or maximum-wait boundary is reached.

    5

    What is a partial batch failure?

    A partial failure occurs when some records succeed while other records in the same batch fail or have uncertain outcomes.

    6

    Should the complete batch be retried?

    Retry the complete batch only when the operation is atomic or every record is independently idempotent. Otherwise, retry failed entries separately.

    7

    What is batch splitting?

    Batch splitting divides a failed batch into smaller groups to identify or isolate the failing record.

    8

    How does batching affect ordering?

    Parallel processing inside a batch can change completion order. Records belonging to one ordered key should be processed sequentially.

    9

    How does Kafka use batches?

    Kafka producers can group records for efficient publication, and consumers can retrieve several records from their assigned partitions in one poll cycle.

    10

    How does RabbitMQ batching work?

    An application can accumulate several deliveries and process them together, while acknowledgments and failures remain aligned with durable message outcomes.

    11

    How does SQS batch processing handle failures?

    Individual entries can require separate success or retry decisions, so the consumer should track each message result independently.

    12

    What is the safest batching strategy?

    Use bounded count, byte, and time limits; per-record idempotency; documented partial-failure handling; ordering-aware processing; bounded concurrency; and durable progress recording.

    Key Takeaway

    Batching improves throughput by sharing fixed network, serialization, transaction, and acknowledgment costs across several records. It also adds waiting latency and increases memory, transaction, retry, and failure scope. Use a hybrid flush policy based on maximum record count, byte size, and waiting time. Keep batches and concurrent batch counts bounded by consumer and downstream capacity. Define whether a batch is atomic or can partially succeed, and preserve a stable identity and result for every independently meaningful record. Kafka batch consumers must manage offsets per partition, RabbitMQ consumers must align acknowledgments with durable outcomes, and SQS consumers must handle individual batch-entry results. Preserve order within the required business key, use idempotent bulk writes, retry only unresolved records where possible, and move permanent failures to a DLQ. Finally, measure both throughput and waiting latency and test partial failures, retries, shutdown, downstream saturation, and recovery.