Table of Contents

    autoscaling

    LOAD BALANCING, PROXIES & ELASTIC SCALING

    Autoscaling

    Learn how autoscaling observes workload signals, calculates required capacity, adds or removes application instances, integrates with load balancers and health checks, handles startup and shutdown safely, prevents scaling oscillation, protects downstream services, and balances performance, availability, and cost.

    Introduction

    Application traffic is rarely constant. An online learning platform can receive ordinary traffic during most of the day and substantially higher traffic during course launches, examinations, live sessions, campaigns, or scheduled learning activities.

    If the deployment always runs with minimum capacity, sudden demand can cause:

    • High response time
    • Request timeouts
    • Queue backlog
    • CPU or memory saturation
    • Connection-pool exhaustion
    • Reduced application availability

    If the deployment always runs with maximum capacity, much of that capacity can remain unused during low-demand periods.

    Autoscaling automatically changes allocated capacity according to observed metrics, schedules, forecasts, or configured policies.

    Core idea: Autoscaling is a feedback-control process. It observes demand, compares that demand with a target, calculates desired capacity, applies a scaling action, and then waits for the system to stabilize before making further decisions.

    Observe workload
          |
          v
    Compare metric with target
          |
          v
    Calculate desired capacity
          |
          +-- Demand increased:
          |      scale out
          |
          +-- Demand decreased:
                 scale in
          |
          v
    Wait for new capacity to stabilize
          |
          v
    Observe again

    Prerequisites

    # Prerequisite Why It Is Needed
    1 Vertical and horizontal scaling Autoscaling commonly changes the number of horizontally scaled instances.
    2 Load balancing New healthy instances must begin receiving traffic, while removed instances must leave routing safely.
    3 Health checks An instance should not receive production traffic before it becomes ready.
    4 Statelessness and sessions Instances must be replaceable without losing required user state.
    5 Metrics and observability Autoscaling decisions depend on accurate workload and performance signals.
    6 Queues and background workers Worker capacity is commonly scaled from backlog and processing-delay signals.
    7 Databases and connection pools Adding application instances can increase pressure on shared dependencies.

    What Is Autoscaling?

    Autoscaling is the automatic adjustment of computing capacity to match changing workload requirements.

    When demand increases, the platform can add resources to maintain the required performance. When demand decreases, unnecessary resources can be removed to reduce cost.

    Low demand:
    
    3 application instances
    
    
    Traffic increases:
    
    3 -> 5 -> 8 instances
    
    
    Traffic decreases:
    
    8 -> 6 -> 4 -> 3 instances

    Autoscaling is most commonly associated with horizontal scaling, where the system adds or removes instances. Vertical resource changes can also be automated in some environments, but resizing can require restart or redeployment and is therefore usually less dynamic.

    Autoscaling Goal
    maintain service objectives → respond to changing demand → retain failure headroom → avoid unnecessary capacity

    Scaling Terms

    Term Meaning
    Scale out Add more instances
    Scale in Remove instances
    Scale up Increase the resources assigned to an existing instance
    Scale down Reduce the resources assigned to an existing instance
    Desired capacity The number of instances the scaling controller currently wants
    Minimum capacity The lowest permitted number of instances
    Maximum capacity The highest permitted number of instances
    Cooldown A stabilization period after a scaling action

    Scalability vs Elasticity vs Autoscaling

    Concept Meaning
    Scalability The system can handle more load by receiving additional capacity.
    Elasticity Capacity can grow and shrink as demand changes.
    Autoscaling Policies and controllers adjust capacity automatically.

    An application can be horizontally scalable but not automatically scaled. Operators can still add and remove instances manually.

    Components of an Autoscaling System

    A complete autoscaling strategy contains several parts:

    1. Instrumentation that captures workload and resource metrics
    2. A monitoring system that collects and aggregates those metrics
    3. A scaling policy containing limits, targets, and rules
    4. A controller that calculates desired capacity
    5. A provisioning system that creates or removes resources
    6. Health checks that determine when new resources are usable
    7. A load balancer that updates backend membership
    8. Observability that records scaling decisions and outcomes
    Application and infrastructure metrics
                     |
                     v
              Monitoring system
                     |
                     v
             Autoscaling policy
                     |
                     v
            Scaling controller
                     |
                     v
         Provision or remove instances
                     |
                     v
           Health and readiness checks
                     |
                     v
              Load-balancer pool

    Scaling Signals

    A scaling signal is a measured value used to decide whether capacity should change.

    Signal Suitable Workload Main Limitation
    CPU utilization CPU-bound request or computation processing Does not directly identify memory, I/O, or dependency bottlenecks
    Memory utilization Workloads whose memory use follows demand predictably Caches and memory leaks can distort the signal
    Requests per second Request costs are relatively predictable Different endpoints can require substantially different work
    Concurrent requests Operations remain active for meaningful periods One expensive request can differ from many inexpensive requests
    Response time User latency reflects capacity pressure Scaling begins after degradation has already appeared
    Queue length Workers process queued background jobs Message processing times can vary
    Oldest-message age Backlog waiting time is the important objective A blocked or poison message can distort the measurement
    Active connections WebSocket, streaming, database, or long-lived TCP workloads Connection activity and resource cost can vary
    Schedule Traffic follows a known recurring pattern Unexpected demand still requires reactive protection

    CPU-based Autoscaling

    CPU utilization is a common autoscaling signal because it is widely available and simple to observe.

    Average CPU remains above
    the approved upper target
          |
          v
    Add application instances
    
    
    Average CPU remains below
    the approved lower target
          |
          v
    Remove application instances cautiously

    CPU-based scaling is useful only when adding instances distributes the CPU-producing work.

    Additional application instances do not correct:

    • A slow database query
    • A shared database lock
    • An exhausted downstream rate limit
    • A memory leak
    • A hot partition
    • A blocked external service

    Memory-based Autoscaling

    Memory can be useful when demand produces predictable per-request or per-worker memory consumption.

    Memory requires careful interpretation because:

    • Runtime memory can remain allocated after traffic decreases
    • Local caches can intentionally consume available memory
    • A leak can cause continuous scale-out without recovery
    • New instances can repeat the same memory problem

    Metric rule: Select a signal that changes because of useful workload demand. Autoscaling should not hide a leak, inefficient query, dependency failure, or defective application.

    Queue-based Autoscaling

    Background workers can scale according to queue backlog.

    Producers
        |
        v
    Durable queue
        |
        +-- Worker 1
        +-- Worker 2
        +-- Worker 3

    When backlog grows, more workers can be added. When the queue remains empty or within the approved lower range, workers can be removed.

    Queue length alone can be misleading. A queue containing 1,000 one-second jobs differs greatly from a queue containing 1,000 ten-minute jobs.

    Backlog-per-worker

    A simplified signal is:

    \[ BacklogPerWorker = \frac{ VisibleMessages }{ ActiveWorkers } \]

    A time-oriented estimate can also consider average processing duration:

    \[ EstimatedDrainTime = \frac{ QueueLength \times AverageProcessingTime }{ ActiveWorkers } \]

    These formulas are planning approximations. Retries, message variation, ordering, dependency limits, and concurrency affect actual performance.

    Main Autoscaling Approaches

    Approach How It Works Best Fit
    Threshold-based Adds or removes capacity when a metric crosses configured limits Simple and predictable workloads
    Target tracking Adjusts capacity to keep a metric near a target value Workload metrics that correlate with required capacity
    Step scaling Changes capacity by different amounts according to overload severity Workloads where small and severe overloads need different responses
    Scheduled scaling Changes capacity at configured times Predictable events or recurring traffic patterns
    Predictive scaling Uses observed history or forecasts to provision capacity before expected demand Workloads with sufficiently repeatable demand patterns

    Threshold-based Scaling

    If CPU remains above
    the upper threshold:
    
    Add capacity
    
    
    If CPU remains below
    the lower threshold:
    
    Remove capacity

    Use separate upper and lower thresholds. This creates a stable range and reduces repeated scale-out and scale-in actions around one boundary.

    Scale-out threshold:
    
    Higher utilization boundary
    
    
    Stable operating range:
    
    No scaling action
    
    
    Scale-in threshold:
    
    Lower utilization boundary

    Target-tracking Scaling

    Target tracking attempts to keep a selected metric near a configured target.

    A simplified desired-capacity estimate is:

    \[ DesiredInstances = CurrentInstances \times \frac{ CurrentMetric }{ TargetMetric } \]

    Example

    Current instances:
    
    4
    
    
    Current average utilization:
    
    80%
    
    
    Target utilization:
    
    50%
    
    
    Estimated desired instances:
    
    4 × 80 / 50
    =
    6.4
    
    
    Controller applies platform-specific
    rounding and policy limits.

    Exact calculations, rounding, stabilization, and evaluation behaviour depend on the autoscaling platform.

    Step Scaling

    Small overload:
    
    Add 1 instance
    
    
    Moderate overload:
    
    Add 3 instances
    
    
    Severe overload:
    
    Add 6 instances

    Step scaling can react more strongly to severe demand but must remain within the configured maximum capacity and downstream limits.

    Scheduled Scaling

    Scheduled scaling changes capacity before or during a predictable event.

    Known examination begins:
    
    10:00
    
    
    Scheduled scale-out begins:
    
    Before traffic arrives
    
    
    At 10:00:
    
    Additional instances are already
    healthy and receiving traffic

    Scheduled scaling is useful because reactive scaling requires time to detect load, provision resources, initialize the application, and pass readiness checks.

    Predictive Scaling

    Predictive scaling attempts to provision capacity before expected workload based on historical or forecast information.

    Historical usage pattern
          |
          v
    Forecast future demand
          |
          v
    Provision capacity in advance
          |
          v
    Validate actual demand
          |
          v
    Reactive rules handle differences

    Forecasts can be wrong. Predictive scaling should operate within minimum, maximum, cost, and safety boundaries and should be complemented by reactive protection.

    Startup Delay

    A scaling action does not create usable capacity immediately.

    Scaling decision
          |
          v
    Provision infrastructure
          |
          v
    Start runtime
          |
          v
    Load application
          |
          v
    Load configuration and secrets
          |
          v
    Initialize dependencies
          |
          v
    Warm required components
          |
          v
    Pass readiness check
          |
          v
    Join load-balancer pool

    The full period from scaling decision to usable capacity is the effective scale-out delay.

    Capacity rule: Autoscaling cannot replace minimum headroom. Existing capacity must continue serving users while additional instances are starting.

    Application Warm-up

    A newly started instance can pass a basic process check before reaching normal production efficiency.

    Warm-up can include:

    • Creating database connections
    • Initializing runtime components
    • Loading application configuration
    • Loading templates or metadata
    • Building safe local caches
    • Establishing downstream clients

    Readiness should prevent premature traffic. Gradual traffic ramp-up can also protect a new instance where supported.

    Cooldown and Stabilization

    A cooldown or stabilization period gives a previous scaling action time to affect the observed workload.

    Scale out
        |
        v
    New instances starting
        |
        v
    Metric remains temporarily high
        |
        v
    Without stabilization:
    
    Controller scales out repeatedly
    
    
    With stabilization:
    
    Controller waits for the previous
    capacity change to take effect

    Cooldown should account for provisioning, startup, readiness, connection establishment, cache warm-up, and metric-collection delay.

    Scaling Oscillation

    Oscillation, also called flapping, occurs when capacity repeatedly scales out and in.

    Load rises
        -> Scale out
    
    Load falls briefly
        -> Scale in
    
    Capacity falls
        -> Load rises
    
    Scale out again

    Possible controls include:

    • Separate scale-out and scale-in thresholds
    • Longer scale-in observation periods
    • Cooldown or stabilization windows
    • Minimum instance lifetimes
    • Smaller scale-in steps
    • Maximum scale-in rate
    • Workload-aware metrics

    Scaling-direction rule: Scale out quickly enough to protect performance, but scale in conservatively enough to avoid removing capacity during a temporary reduction in demand.

    Minimum Capacity

    Minimum capacity is the lowest instance count the autoscaling system can maintain.

    It should consider:

    • Ordinary traffic
    • Expected instance failure
    • Deployment activity
    • Startup delay
    • Availability-zone or failure-domain requirements
    • Maintenance operations
    • Sudden traffic before scaling completes

    Minimum capacity should not be selected only from the lowest observed traffic.

    Maximum Capacity

    Maximum capacity prevents uncontrolled infrastructure growth.

    It should consider:

    • Cost boundaries
    • Database connection capacity
    • Downstream API limits
    • Queue and broker capacity
    • Network and load-balancer limits
    • Available address space
    • Service quotas
    • Licensing constraints

    Reaching maximum capacity should generate a clear alert because additional demand can no longer be addressed by scaling the configured resource group.

    Downstream Bottlenecks

    Scaling one tier increases demand on its dependencies.

    Before scaling:
    
    4 API instances
    20 database connections each
    =
    80 potential connections
    
    
    After scaling:
    
    20 API instances
    20 database connections each
    =
    400 potential connections

    The API tier gained capacity, but the database can now face connection, query, CPU, memory, lock, and storage pressure.

    Connection Budget

    A simplified connection estimate is:

    \[ MaximumConnections = MaximumInstances \times PoolSizePerInstance \]

    Reserve capacity for workers, reporting, administration, deployment, monitoring, and recovery activity.

    Hot Partitions and Autoscaling

    Autoscaling does not automatically fix uneven workload distribution.

    10 application instances
          |
          v
    All requests target
    one database partition
          |
          v
    Hot partition remains
    the system bottleneck

    Add instances only when the workload can be distributed. Hot keys, serialized operations, global counters, and locked resources can remain bottlenecks after compute scale-out.

    Safe Scale-in

    Removing capacity requires graceful shutdown.

    Select instance for scale-in
          |
          v
    Mark readiness as false
          |
          v
    Load balancer stops new requests
          |
          v
    Drain in-flight requests
          |
          v
    Stop message consumption
          |
          v
    Complete or release owned work
          |
          v
    Close connections
          |
          v
    Terminate instance

    Scale-in policy must consider:

    • Long-running requests
    • Uploads
    • WebSocket connections
    • Background jobs
    • Queue visibility timeouts
    • Local temporary processing
    • Connection-drain duration

    Autoscaling and Sessions

    Required sessions should not exist only in one instance's memory.

    Scale-in removes Instance A
          |
          v
    User's next request reaches Instance C
          |
          v
    Instance C retrieves the session
    from a shared store
          |
          v
    User workflow continues

    Sticky sessions complicate scale-in because active users can remain bound to an instance selected for removal.

    Autoscaling and Health Checks

    New instances should enter the load-balancer pool only after readiness succeeds.

    New instance created
          |
          v
    Startup check
          |
          v
    Application initialization
          |
          v
    Readiness succeeds
          |
          v
    Load balancer adds backend

    An instance that cannot initialize should not count as usable capacity merely because it exists.

    Failure Headroom

    Capacity planning should account for expected failures.

    Required normal capacity:
    
    6 instances
    
    
    One instance unavailable:
    
    5 healthy instances remain
    
    
    Question:
    
    Can 5 instances safely handle
    the required workload while a
    replacement is starting?

    Running every instance close to saturation leaves little capacity for failure, deployment, or sudden traffic.

    Load Shedding

    Autoscaling is not instantaneous. During severe demand, the system can still require overload protection.

    Possible controls include:

    • Rate limiting
    • Concurrency limiting
    • Bounded queues
    • Admission control
    • Serving approved stale cached data
    • Deferring nonessential work
    • Reducing optional response details
    • Returning explicit overload responses

    Do not silently degrade correctness-critical operations such as payment, inventory reservation, authorization changes, or financial posting.

    Retry Amplification

    Application becomes slow
          |
          v
    Clients time out
          |
          v
    Clients retry
          |
          v
    Traffic increases
          |
          v
    Autoscaling detects additional load
          |
          v
    New capacity starts too late
    or downstream system becomes overloaded

    Use:

    • Bounded retry attempts
    • Exponential backoff
    • Randomized jitter
    • End-to-end deadlines
    • Idempotency
    • Overload-aware admission control

    Web-tier Example

    Client traffic
          |
          v
    Application Load Balancer
          |
          +-- Web Instance 1
          +-- Web Instance 2
          +-- Web Instance 3
                  |
                  v
            Shared services
    
    
    Autoscaling signal:
    
    Requests per healthy instance
    
    
    Scale out:
    
    Add web instances
    
    
    Scale in:
    
    Drain and remove web instances

    The application instances should be stateless, health checked, and protected by minimum and maximum capacity.

    Worker-tier Example

    Course-video uploads
          |
          v
    Processing queue
          |
          +-- Worker 1
          +-- Worker 2
          +-- Worker 3
    
    
    Scaling signals:
    
    - Visible messages
    - Oldest-message age
    - Processing duration
    - Downstream capacity

    Worker scale-out should not exceed the safe concurrency of object storage, media-processing dependencies, databases, or third-party APIs.

    Conceptual Autoscaling Policy

    autoscaling:
      minimumInstances: 3
      maximumInstances: 20
      defaultInstances: 4
    
      metric:
        name: requests-per-healthy-instance
        target: approved-target
    
      scaleOut:
        evaluationWindow: approved-short-window
        maximumStep: approved-scale-out-step
    
      scaleIn:
        evaluationWindow: approved-long-window
        maximumStep: approved-scale-in-step
    
      stabilization:
        afterScaleOut: approved-scale-out-cooldown
        afterScaleIn: approved-scale-in-cooldown
    
      readiness:
        path: /health/ready
    
      termination:
        markNotReady: true
        drainRequests: true
        gracePeriod: approved-drain-period

    This is conceptual configuration. Metric names, thresholds, timing, and policy structures depend on the selected platform and tested workload.

    Kubernetes Horizontal Pod Autoscaling

    Kubernetes can use a Horizontal Pod Autoscaler to adjust the desired replica count of a supported workload according to observed resource or custom metrics.

    Conceptual HPA Definition

    apiVersion: autoscaling/v2
    kind: HorizontalPodAutoscaler
    metadata:
      name: course-api
    
    spec:
      scaleTargetRef:
        apiVersion: apps/v1
        kind: Deployment
        name: course-api
    
      minReplicas: 3
      maxReplicas: 20
    
      metrics:
        - type: Resource
          resource:
            name: cpu
            target:
              type: Utilization
              averageUtilization: approved-target

    Resource requests, metrics availability, readiness, application startup, and node capacity must also be configured correctly. Adding pods does not help when there is no cluster capacity on which to schedule them.

    Pod Scaling vs Node Scaling

    Layer Scaling Action Important Dependency
    Application or pod layer Add or remove application replicas Enough compute capacity must already exist
    Node or cluster layer Add or remove worker machines Infrastructure provisioning and scheduling
    Database layer Resize, replicate, or partition according to platform Consistency, routing, storage, and recovery

    These layers can have different startup times and policies. Application replicas can remain pending while new compute nodes are being created.

    Database Autoscaling

    Database scaling differs from adding stateless application instances.

    Possible database scaling actions include:

    • Increase compute or memory
    • Add read replicas
    • Increase provisioned throughput
    • Expand storage
    • Rebalance or partition data

    Review transaction behaviour, replication lag, connection routing, consistency, scale-in safety, and recovery before automating a database capacity change.

    Cost Considerations

    Autoscaling can reduce unused capacity, but it does not guarantee minimum cost.

    A simplified compute-cost model is:

    \[ ComputeCost = \sum \left( InstanceCount \times InstancePrice \times RunningDuration \right) \]

    Also consider:

    • Minimum always-on capacity
    • Load-balancer cost
    • Metrics and logging cost
    • Data-transfer cost
    • Database and cache scaling
    • Startup and warm-up inefficiency
    • Rapid scaling oscillation
    • Reserved or committed capacity
    • Licensing

    Security Considerations

    Automatically created instances must receive the same approved security configuration as existing instances.

    Ensure:

    • Instances are created from approved images or artifacts
    • Secrets are retrieved through approved mechanisms
    • Network restrictions apply automatically
    • Identity and permissions use least privilege
    • Monitoring and security tooling initialize correctly
    • Temporary instances do not retain sensitive local state
    • Scale-in follows secure cleanup and lifecycle policies

    Observability

    Useful autoscaling metrics include:

    • Current instance count
    • Desired instance count
    • Healthy and ready instance count
    • Scaling actions by direction
    • Scaling-action failures
    • Instance provisioning time
    • Application startup time
    • Readiness delay
    • CPU and memory per instance
    • Requests per instance
    • Queue backlog per worker
    • Oldest-message age
    • Connection-pool usage
    • Scale-in drain duration
    • Time spent at maximum capacity
    • Cost by service and scaling group

    Scaling-event Record

    {
      "service": "course-api",
      "action": "scale-out",
      "previousCapacity": 4,
      "desiredCapacity": 6,
      "metric": "requests-per-healthy-instance",
      "reason": "target exceeded",
      "policyVersion": "approved-policy-version"
    }

    Scaling records should help operators understand why capacity changed without including credentials or sensitive request data.

    Alert Conditions

    Alert when:

    • The service reaches maximum capacity
    • Desired capacity cannot be provisioned
    • New instances repeatedly fail readiness
    • Scale-out does not reduce latency or backlog
    • Autoscaling oscillates repeatedly
    • Healthy capacity falls below the required minimum
    • Database connections approach their safe limit
    • Queue backlog grows after worker scale-out
    • Scale-in interrupts active work
    • One failure domain loses too much capacity
    • Scaling cost grows unexpectedly
    • Metrics required by the scaling controller are unavailable

    Troubleshooting Workflow

    1. Identify the user-visible performance or availability symptom.
    2. Confirm the current and desired capacity.
    3. Identify which scaling policy is active.
    4. Inspect the metric that triggered or failed to trigger scaling.
    5. Check metric delay, aggregation, and missing data.
    6. Check minimum, maximum, and default capacity.
    7. Check cooldown and stabilization behaviour.
    8. Check provisioning failures and service quotas.
    9. Check application startup and readiness.
    10. Check load-balancer backend membership.
    11. Check database, cache, queue, and external dependency capacity.
    12. Check retry amplification and overload controls.
    13. Check graceful scale-in and connection draining.
    14. Compare the result with representative load-test evidence.

    Common Autoscaling Mistakes

    1

    Selecting the Wrong Metric

    Capacity changes without correcting the actual user-facing or processing bottleneck.

    2

    Using CPU as the Only Signal

    Database latency, queue age, memory, connections, or storage can be saturated while CPU remains moderate.

    3

    Scaling Out Too Late

    Traffic exceeds capacity before new instances finish provisioning and startup.

    4

    Scaling In Too Aggressively

    Capacity is removed during a temporary reduction and must be recreated immediately afterward.

    5

    Using the Same Boundary for Scale-out and Scale-in

    Minor metric variation can cause repeated scaling oscillation.

    6

    Ignoring Startup and Warm-up Time

    Newly created resources exist but do not provide usable capacity soon enough.

    7

    Counting Unready Instances as Capacity

    The scaling controller believes adequate capacity exists while the load balancer has too few usable backends.

    8

    Ignoring Downstream Limits

    Additional instances overwhelm database connections, queues, caches, or external APIs.

    9

    Using Autoscaling to Hide a Memory Leak

    The application repeatedly adds instances without correcting the defect.

    10

    Removing Instances without Draining

    Scale-in interrupts requests, uploads, messages, and long-lived connections.

    11

    Setting Maximum Capacity without an Alert

    The service stops scaling while operators remain unaware that no additional capacity is available.

    12

    Testing Only Uniform Traffic

    The policy remains untested against bursts, hot keys, expensive routes, retries, cache misses, and dependency failures.

    Recommended Test Cases

    Test Expected Evidence
    Normal scale-out New instances become ready and receive traffic
    Normal scale-in Instances drain before termination
    Sudden traffic burst Minimum headroom protects users while new capacity starts
    Sustained traffic increase Capacity converges to the required level
    Temporary metric spike The policy avoids unnecessary repeated scaling
    Traffic reduction Scale-in occurs only after the approved stabilization period
    Maximum capacity The platform raises an alert and applies overload protection
    Instance startup failure The failed instance receives no production traffic
    Readiness delay New capacity is counted only after becoming usable
    One instance failure Remaining capacity serves traffic while replacement occurs
    Database connection limit Maximum application scale does not exhaust the database budget
    Queue backlog Worker scale-out reduces processing delay safely
    Poison message One unprocessable message does not trigger uncontrolled scaling
    Retry storm Backoff, jitter, and admission control protect capacity
    Long-lived connections Scale-in drains or reconnects clients according to policy
    Autoscaling oscillation Threshold separation and stabilization prevent flapping

    Autoscaling Best Practices

    Recommended Practices

    • Confirm that the workload can scale horizontally.
    • Keep scalable application instances stateless and replaceable.
    • Select metrics that correlate with useful workload demand.
    • Use queue delay or backlog for asynchronous workers.
    • Maintain meaningful minimum and maximum capacity.
    • Retain headroom for startup delay and expected failures.
    • Use separate scale-out and scale-in conditions.
    • Scale out faster than scale in when availability is the priority.
    • Use cooldown and stabilization windows.
    • Count only healthy and ready instances as usable capacity.
    • Warm new instances before sending significant traffic.
    • Budget database and downstream connections at maximum scale.
    • Limit worker concurrency according to downstream capacity.
    • Use scheduled scaling for known demand where appropriate.
    • Keep reactive rules for unexpected demand.
    • Drain requests, jobs, and connections before scale-in.
    • Use rate limiting and load shedding while capacity starts.
    • Alert when maximum capacity is reached.
    • Record every scaling decision and policy version.
    • Load-test scale-out, scale-in, failures, and downstream saturation.

    Practice Exercise

    Design autoscaling for your online learning platform's API and video-processing workers.

    Requirements

    1. Measure the safe capacity of one API instance.
    2. Estimate normal and peak request rates.
    3. Select minimum and maximum API capacity.
    4. Choose a workload metric for API scale-out.
    5. Define separate scale-out and scale-in conditions.
    6. Measure application startup and readiness time.
    7. Maintain capacity while new instances initialize.
    8. Set a maximum database-connection budget.
    9. Create a worker autoscaling policy based on queue backlog and age.
    10. Limit workers according to media-processing dependency capacity.
    11. Add graceful request and message draining.
    12. Add rate limiting for severe API demand.
    13. Test a sudden examination-period traffic burst.
    14. Test a poison message in the processing queue.
    15. Test one failed instance during peak load.
    16. Test the maximum-capacity scenario.
    17. Measure performance, failure behaviour, and cost.

    Autoscaling-design Template

    Workload Scaling Signal Capacity Boundary Primary Safety Control
    Web and API requests Requests per healthy instance and latency Minimum and maximum ready instances Readiness, connection budget, and rate limiting
    Video-processing workers Queue backlog and oldest-message age Maximum safe processing concurrency Durable queue, idempotency, and dependency protection
    WebSocket service Active connections and connection growth Connection capacity per instance Reconnection and graceful draining
    Database readers Read load, connection use, and replica capacity Platform and consistency limits Replication-lag and routing controls
    Scheduled examination traffic Known event schedule plus reactive metrics Forecast capacity and cost boundary Pre-scaling and maximum-capacity alerts

    Frequently Asked Questions

    1

    What is autoscaling?

    Autoscaling automatically adjusts resource capacity according to observed metrics, schedules, forecasts, and configured limits.

    2

    What is scale out?

    Scale out means adding more instances to distribute workload.

    3

    What is scale in?

    Scale in means removing instances when less capacity is required.

    4

    Which metric should autoscaling use?

    Use a metric that reliably reflects the workload and capacity pressure of the specific component, such as CPU, requests per instance, queue age, backlog, or active connections.

    5

    Why is minimum capacity required?

    Minimum capacity serves ordinary traffic and provides headroom while new resources are provisioned or existing instances fail.

    6

    Why is maximum capacity required?

    Maximum capacity limits cost and protects databases, external services, network resources, and other dependencies from uncontrolled concurrency.

    7

    What is a cooldown period?

    A cooldown gives a previous scaling action time to affect the workload before another scaling decision is made.

    8

    What is scaling oscillation?

    Scaling oscillation is repeated scale-out and scale-in caused by unstable thresholds, delayed metrics, or insufficient stabilization.

    9

    Can autoscaling fix a slow database query?

    No. Additional application instances can increase database pressure without correcting the inefficient query.

    10

    Should autoscaling use CPU only?

    Not necessarily. CPU is useful for CPU-bound workloads, but user latency, requests, memory, queue delay, connections, or custom metrics can better represent other workloads.

    11

    How should instances be removed?

    Mark the instance not ready, stop new work, drain existing requests and messages, close connections, and then terminate it.

    12

    Does autoscaling guarantee availability?

    No. The system still requires failure headroom, healthy routing, resilient dependencies, overload protection, tested startup, and safe scale-in.

    Key Takeaway

    Autoscaling automatically adjusts capacity to match changing demand. A reliable autoscaling design begins with a horizontally scalable workload and uses metrics that reflect actual capacity pressure. It defines minimum, maximum, and desired capacity; integrates with readiness checks and load balancing; accounts for startup and warm-up time; and uses separate scale-out and scale-in behaviour to prevent oscillation. Scale out quickly enough to protect users, but scale in cautiously and drain work before termination. Autoscaling cannot correct inefficient queries, hot partitions, memory leaks, or exhausted downstream services. Budget database connections and external concurrency at maximum scale, maintain failure headroom, apply rate limiting while additional capacity starts, and alert when scaling reaches its configured limit. Finally, validate the complete control loop under realistic traffic, failures, deployments, and dependency constraints.