Table of Contents

    Vertical vs horizontal scaling

    LOAD BALANCING, PROXIES & ELASTIC SCALING

    Vertical vs Horizontal Scaling

    Learn how vertical scaling increases the capacity of an existing server, how horizontal scaling adds more server instances, and how statelessness, load balancing, health checks, shared state, databases, queues, autoscaling, graceful shutdown, observability, and cost influence the correct scaling strategy.

    Introduction

    An application that performs well for a small number of users can slow down as traffic, stored data, background work, and concurrent requests increase.

    The system can experience:

    • Higher response time
    • Increased CPU utilization
    • Memory exhaustion
    • Connection-pool saturation
    • Longer queue processing time
    • Database contention
    • More timeouts and retries
    • Reduced availability during failures

    Scaling increases or reduces system capacity so the application can continue meeting its performance, reliability, and cost objectives under changing demand.

    There are two primary scaling directions:

    • Vertical scaling: Increase or decrease the capacity of an existing resource.
    • Horizontal scaling: Add or remove resource instances.

    Core idea: Vertical scaling makes one instance stronger. Horizontal scaling distributes work across several instances. Vertical scaling is usually simpler, while horizontal scaling can provide greater elasticity and resilience when the application is designed for distribution.

    This lesson begins the Load Balancing, Proxies & Elastic Scaling module. The concepts introduced here provide the foundation for load balancers, reverse proxies, health checks, session management, redundancy, failover, and autoscaling.

    Prerequisites

    # Prerequisite Why It Is Needed
    1 CPU, memory, disk, and network fundamentals Scaling decisions depend on the resource that limits the workload.
    2 Processes and threads Application concurrency affects how one instance uses available resources.
    3 HTTP request lifecycle Horizontal scaling distributes incoming requests across several backends.
    4 Databases and connection pools Adding application instances can increase pressure on shared databases.
    5 Caching and sessions Instance-local state can prevent requests from moving safely between servers.
    6 Observability Metrics are required to identify bottlenecks and validate scaling behaviour.

    What Is Scalability?

    Scalability is the ability of a system to maintain acceptable behaviour when workload changes.

    The workload can include:

    • Concurrent users
    • Requests per second
    • Background jobs
    • Messages waiting in queues
    • Database queries
    • Stored records
    • Uploaded files
    • Network traffic

    A scalable system should have an understood way to increase capacity before a bottleneck causes unacceptable latency or failures.

    Scaling Workflow
    measure workload → identify bottleneck → select scaling direction → validate under load → monitor cost and reliability

    Scalability vs Elasticity

    Concept Meaning
    Scalability The system can handle additional workload by increasing resources.
    Elasticity The system can add and remove resources in response to changing demand.
    Capacity The amount of workload the current deployment can handle within its objectives.
    Efficiency The useful work produced from the allocated resources and cost.

    A system can be scalable without being automatically elastic. For example, operators can manually increase server capacity when growth occurs.

    What Is Vertical Scaling?

    Vertical scaling changes the capacity of an existing resource while keeping the number of instances conceptually unchanged.

    Before vertical scaling:
    
    1 application server
    2 CPU cores
    8 GB memory
    
    
    After scaling up:
    
    1 application server
    8 CPU cores
    32 GB memory

    Increasing capacity is called scaling up. Reducing capacity is called scaling down.

    Resources That Can Be Increased

    • CPU cores
    • Processor performance
    • Memory
    • Disk size
    • Disk performance
    • Network capacity
    • Database compute tier
    • Application-service plan size

    Vertical-scaling Architecture

    Clients
       |
       v
    Application Server
    
    Before:
    +----------------------+
    | 2 CPU                |
    | 8 GB RAM             |
    | Limited disk IOPS    |
    +----------------------+
    
    
    Scale up
       |
       v
    
    After:
    +----------------------+
    | 8 CPU                |
    | 32 GB RAM            |
    | Higher disk IOPS     |
    +----------------------+

    The application architecture remains mostly unchanged. The existing instance receives more capacity.

    Advantages of Vertical Scaling

    • Usually simpler than distributing the application
    • Requires fewer instances to configure and monitor
    • Can work for stateful or legacy applications
    • Can avoid distributed session and coordination concerns
    • Can improve single-process or single-thread performance
    • Can be an effective short-term response to gradual growth
    • Can require fewer application-code changes

    Limitations of Vertical Scaling

    • Every resource type has an upper hardware or service-tier limit
    • Scaling can require a restart, redeployment, or temporary unavailability
    • One instance can remain a single failure point
    • Larger instances can be expensive
    • Unused capacity can remain allocated during low traffic
    • One process might not use all added CPU cores effectively
    • Scaling one component does not fix a bottleneck in another component

    Vertical-scaling rule: Scaling up can provide immediate capacity, but it does not remove the maximum-instance limit or automatically improve availability.

    What Is Horizontal Scaling?

    Horizontal scaling changes capacity by adding or removing instances.

    Before horizontal scaling:
    
    1 application instance
    
    
    After scaling out:
    
    Application instance 1
    Application instance 2
    Application instance 3
    Application instance 4

    Adding instances is called scaling out. Removing instances is called scaling in.

    Horizontal-scaling Architecture

                         +------------------+
    Clients ------------>|  Load Balancer   |
                         +------------------+
                               |  |  |
                 +-------------+  |  +-------------+
                 |                |                |
                 v                v                v
    
          +-------------+  +-------------+  +-------------+
          | App         |  | App         |  | App         |
          | Instance 1  |  | Instance 2  |  | Instance 3  |
          +-------------+  +-------------+  +-------------+

    A load balancer, gateway, proxy, service-discovery mechanism, or messaging system distributes work among the available instances.

    Advantages of Horizontal Scaling

    • Capacity can grow by adding instances
    • Unhealthy instances can be removed from service
    • Traffic can be distributed across several backends
    • Instances can be added or removed elastically
    • Rolling deployments become practical
    • The system can tolerate selected instance failures
    • Smaller commodity instances can replace one very large server
    • Different regions or availability zones can host instances

    Challenges of Horizontal Scaling

    • The application must work correctly across several instances
    • Instance-local sessions can cause routing problems
    • Shared caches and databases can become bottlenecks
    • Logs and traces are distributed
    • Deployments require version compatibility
    • Requests can arrive at different instances
    • Background jobs can accidentally run more than once
    • Coordination and distributed consistency become more complex
    • Scale-in must not terminate instances with unfinished work

    Vertical vs Horizontal Scaling

    Area Vertical Scaling Horizontal Scaling
    Method Increase the capacity of an existing instance Add more instances
    Common terms Scale up and scale down Scale out and scale in
    Application changes Often fewer changes Can require statelessness and distributed coordination
    Upper limit Limited by the largest available resource size Limited by architecture, coordination, shared dependencies, and service limits
    Availability One instance can remain a failure point Several healthy instances can provide redundancy
    Load distribution Work remains on one larger instance Work is distributed across instances
    Scaling interruption Can require restart or redeployment New instances can often join while existing instances continue serving
    Operational complexity Usually lower Usually higher
    Elasticity Less commonly automated Commonly automated in cloud environments
    Suitable workload Stateful, legacy, or single-instance-oriented workload Stateless or distributable workload

    Capacity Model

    If one healthy application instance safely processes \(C\) requests per second and the expected peak workload is \(Q\), a simplified instance-count estimate is:

    \[ RequiredInstances = \left\lceil \frac{Q}{C} \right\rceil \]

    Additional capacity should be considered for failures, deployments, unexpected bursts, and measurement uncertainty.

    Example

    Tested safe capacity per instance:
    
    500 requests per second
    
    
    Expected peak traffic:
    
    1,800 requests per second
    
    
    Minimum calculated instances:
    
    ceil(1800 / 500)
    =
    4 instances

    This simplified estimate does not include failure headroom, uneven request cost, downstream limits, or high-percentile latency.

    Failure Headroom

    A horizontally scaled service should remain within acceptable performance after losing an expected number of instances.

    Normal deployment:
    
    4 instances
    
    
    Failure scenario:
    
    1 instance unavailable
    
    
    Remaining capacity:
    
    3 instances
    
    
    Question:
    
    Can 3 instances safely handle
    the required workload?

    Capacity planning should test the degraded scenario rather than assuming every instance is always available.

    Load Balancer

    A load balancer receives requests and distributes them across eligible backend instances.

    Client request
          |
          v
    Load balancer
          |
          +-- Backend 1
          +-- Backend 2
          +-- Backend 3

    A load balancer can consider:

    • Backend health
    • Configured routing algorithm
    • Connection count
    • Backend capacity or weight
    • Geographic location
    • Protocol and request attributes

    Exact features depend on the selected load balancer.

    Basic Load-balancing Algorithms

    Algorithm General Behaviour
    Round robin Distributes requests across backends in sequence
    Weighted round robin Sends more requests to backends assigned greater weight
    Least connections Prefers the backend with fewer active connections
    Weighted least connections Considers both active connections and backend weight
    Hash-based routing Maps a selected request value consistently to a backend
    Resource-aware routing Uses available backend health or utilization information

    The best algorithm depends on request duration, instance capacity, connection behaviour, and the availability of reliable backend metrics.

    Health Checks

    A load balancer should route traffic only to backends considered healthy.

    Load balancer
          |
          v
    Health-check request
          |
          +-- Healthy response:
          |      backend remains eligible
          |
          +-- Failed response:
                 backend is removed
                 from active rotation

    Basic Health Endpoint

    GET /health/live HTTP/1.1
    Host: app.example.com
    {
      "status": "healthy"
    }

    A liveness check indicates whether the process is running. A readiness check indicates whether the instance is prepared to receive traffic.

    Liveness vs Readiness

    Check Question Possible Action
    Liveness Is the process alive and capable of making progress? Restart an unhealthy process
    Readiness Can the instance safely receive new requests? Add or remove the instance from traffic rotation
    Startup Has initialization completed? Delay other health decisions during startup

    Health-check rule: A process can be alive but not ready. Do not send production traffic to an instance until configuration, dependencies, migrations, and initialization required for request handling are complete.

    Avoid Fragile Health Checks

    A health endpoint should not become slow or unstable because it performs large queries or checks every downstream system.

    Fragile check
    Health request performs:
    
    - Large database query
    - External payment API call
    - Search query
    - Object-storage download
    - Cache write
    Focused check
    Liveness:
    
    Process can make progress
    
    
    Readiness:
    
    Instance has required configuration
    and can safely serve its core request path

    Stateless Application Instances

    Horizontal scaling is easier when any healthy instance can process the next request.

    Request 1
        -> Instance A
    
    
    Request 2
        -> Instance C
    
    
    Request 3
        -> Instance B
    
    
    All requests remain correct.

    This does not mean the complete application has no state. It means request-critical state is stored outside one specific application process or is included safely in the request.

    The In-memory Session Problem

    Login request
          |
          v
    Instance A stores session in memory
    
    
    Next request
          |
          v
    Load balancer selects Instance B
    
    
    Instance B:
    
    Session not found

    Possible solutions include:

    • Store sessions in a shared key-value store
    • Use appropriately protected self-contained tokens
    • Use sticky sessions as a controlled transitional strategy
    • Store authoritative session state in a shared database

    Sticky Sessions

    Sticky sessions attempt to route one client to the same backend instance.

    User A
        -> Instance 1
        -> Instance 1
        -> Instance 1
    
    
    User B
        -> Instance 2
        -> Instance 2
        -> Instance 2

    Sticky routing can simplify migration from instance-local sessions, but it can cause uneven distribution and does not protect state when the selected instance fails.

    Session rule: Sticky sessions can reduce routing flexibility. Externalizing important session state usually provides stronger horizontal-scaling and failover behaviour.

    The Local-file Problem

    Files written to one application instance might not be available on another instance.

    Upload handled by Instance A
    
    File stored at:
    
    /var/app/uploads/lesson.pdf
    
    
    Download handled by Instance B
    
    Result:
    
    File not found

    Shared file storage, object storage, or another durable shared content system should hold files that must be accessible from multiple instances.

    Shared Database Bottleneck

    Adding application instances can shift the bottleneck to the database.

    Before:
    
    2 application instances
    20 database connections each
    =
    40 possible connections
    
    
    After scaling:
    
    20 application instances
    20 database connections each
    =
    400 possible connections

    The application tier gained capacity, but the database can now face higher connection count, query concurrency, locking, CPU, and storage pressure.

    Connection Budget

    A simplified maximum connection estimate is:

    \[ MaximumConnections = InstanceCount \times PoolSizePerInstance \]

    Reserve capacity for administration, migrations, background jobs, and other database clients.

    Connection-pool Management

    Horizontal scaling should coordinate instance count and per-instance pool size.

    Database connection budget:
    
    200 connections
    
    
    Maximum application instances:
    
    10
    
    
    Initial pool direction:
    
    Less than 20 connections per instance
    
    because capacity is also needed for:
    
    - Background workers
    - Administrative access
    - Deployment operations
    - Monitoring
    - Failure headroom

    Exact pool settings require workload measurement and database-specific guidance.

    Multiple Scaling Layers

    A production system contains several independently scalable components.

    Client traffic
          |
          v
    Gateway or load balancer
          |
          v
    Web/API instances
          |
          v
    Cache
          |
          v
    Database
          |
          v
    Object storage
    
    
    Background flow:
    
    Queue
      |
      v
    Worker instances

    Scaling the web tier does not automatically scale the cache, database, message broker, worker tier, or external dependencies.

    Scaling Databases

    Databases can be scaled vertically and horizontally, but horizontal database scaling is more complex than adding stateless application instances.

    Vertical Database Scaling

    • Increase CPU and memory
    • Increase storage performance
    • Increase service tier
    • Increase connection or throughput capacity where supported

    Horizontal Database Scaling

    • Add read replicas
    • Partition or shard data
    • Separate workloads
    • Create query-specific stores
    • Use distributed database capabilities

    Horizontal database scaling introduces replication, routing, consistency, partitioning, rebalancing, and cross-partition query considerations.

    Read Replicas

                        +------------------+
    Writes ------------>| Primary Database |
                        +------------------+
                               |
                               | Replication
                               v
                 +-------------+-------------+
                 |                           |
                 v                           v
    
          +---------------+           +---------------+
          | Read Replica 1|           | Read Replica 2|
          +---------------+           +---------------+

    Read replicas can distribute suitable read traffic. Replication delay means a replica can temporarily return an older value.

    Reads requiring immediate visibility after a write might need the authoritative primary or another supported consistency mechanism.

    Queue-based Horizontal Scaling

    Background processing can scale by adding consumers to a queue.

    Producers
        |
        v
    Message queue
        |
        +-- Worker 1
        +-- Worker 2
        +-- Worker 3
        +-- Worker 4

    The queue buffers work and consumers process available messages.

    Worker scaling should consider:

    • Queue length
    • Oldest-message age
    • Message processing time
    • Downstream capacity
    • Retry traffic
    • Partition ordering
    • Idempotency

    Autoscaling

    Autoscaling automatically changes resource capacity according to monitored conditions or schedules.

    Metrics
       |
       v
    Autoscaling decision
       |
       +-- Demand increased:
       |      add instances
       |
       +-- Demand decreased:
              remove instances

    Autoscaling commonly applies to horizontal compute scaling. Vertical resizing can be harder to automate because it can require restarting or redeploying a resource.

    Autoscaling Signals

    Signal Useful When Limitation
    CPU utilization The workload is CPU-bound Can miss memory, I/O, or dependency bottlenecks
    Memory utilization The workload consumes memory predictably Can react poorly to leaks or cache behaviour
    Request rate Request cost is reasonably stable Different requests can require different work
    Response time User latency reflects capacity pressure Scaling can be too late after latency rises
    Queue length Workers process independent messages Message duration can vary
    Oldest-message age Backlog delay is the important objective Poison messages can distort the signal
    Schedule Demand follows a predictable calendar pattern Unexpected traffic still requires another response

    Reactive vs Scheduled Scaling

    Reactive Scaling

    Load increases
          |
          v
    Metric crosses threshold
          |
          v
    Scaling action starts
          |
          v
    New instance initializes
          |
          v
    Health check succeeds
          |
          v
    Instance receives traffic

    Reactive scaling has a delay between detecting demand and receiving usable capacity.

    Scheduled Scaling

    Known examination period begins at 10:00
    
    Scale out before 10:00
          |
          v
    Instances become ready
          |
          v
    Traffic arrives

    Scheduled scaling is useful for predictable patterns and can be combined with metric-based scaling.

    Instance Startup Time

    Scaling decisions must account for how long a new instance requires before it becomes ready.

    Instance provisioning
          |
          v
    Runtime startup
          |
          v
    Configuration loading
          |
          v
    Dependency initialization
          |
          v
    Application warm-up
          |
          v
    Readiness success
          |
          v
    Traffic begins

    Slow startup can cause autoscaling to react too late. Maintain enough minimum capacity to absorb traffic while new instances initialize.

    Warm-up Behaviour

    A newly started instance can be healthy but not yet operating at full efficiency.

    Warm-up can include:

    • Loading application code
    • Building runtime caches
    • Opening database connections
    • Loading configuration and secrets
    • Compiling templates
    • Initializing dependency clients

    Gradual traffic ramp-up can protect a new instance from receiving too much work immediately.

    Scaling Oscillation

    Poorly tuned rules can repeatedly add and remove instances.

    CPU rises
      -> scale out
    
    CPU falls
      -> scale in
    
    Traffic rises again
      -> scale out
    
    Traffic falls briefly
      -> scale in

    This behaviour is sometimes called oscillation or flapping.

    Possible controls include:

    • Separate scale-out and scale-in thresholds
    • Cooldown periods
    • Minimum instance lifetime
    • Longer scale-in observation windows
    • Minimum and maximum capacity
    • Step scaling based on overload severity

    Autoscaling rule: Scale out quickly enough to protect users, but scale in cautiously enough to avoid removing capacity during a temporary drop.

    Graceful Scale-in

    Removing an instance safely requires more than terminating the process.

    Select instance for removal
          |
          v
    Mark instance not ready
          |
          v
    Stop assigning new requests
          |
          v
    Allow in-flight work to finish
          |
          v
    Stop background consumers
          |
          v
    Close resources
          |
          v
    Terminate instance

    Long-running requests, uploads, WebSocket connections, and background jobs require explicit draining behaviour.

    Long-lived Connections

    WebSockets and other long-lived connections can complicate load balancing and scale-in.

    Client
        |
        v
    Long-lived connection
        |
        v
    Instance A
    
    
    Scale-in selects Instance A
          |
          v
    Connection needs:
    
    - Graceful closure
    - Reconnection guidance
    - State recovery
    - Updated service discovery

    Connection count can be a more useful scaling metric than simple request rate for long-lived protocols.

    Conceptual Autoscaling Policy

    autoscaling:
      minimumInstances: 3
      maximumInstances: 20
    
      scaleOut:
        metric: request-utilization
        threshold: approved-high-threshold
        evaluationWindow: approved-short-window
        addInstances: 2
    
      scaleIn:
        metric: request-utilization
        threshold: approved-low-threshold
        evaluationWindow: approved-long-window
        removeInstances: 1
    
      cooldown:
        afterScaleOut: approved-cooldown
        afterScaleIn: approved-cooldown
    
      readiness:
        endpoint: /health/ready
    
      termination:
        drainRequests: true
        gracePeriod: approved-drain-period

    This is conceptual configuration. Thresholds and timing values must be established using representative workload tests and platform-specific guidance.

    PHP Readiness Endpoint

    <?php
    
    declare(strict_types=1);
    
    final class ReadinessController
    {
        public function __construct(
            private ApplicationState $applicationState
        ) {
        }
    
        public function check(): void
        {
            header(
                'Content-Type: application/json'
            );
    
            if (!$this->applicationState->isReady()) {
                http_response_code(
                    503
                );
    
                echo json_encode([
                    'status' => 'not-ready'
                ]);
    
                return;
            }
    
            http_response_code(
                200
            );
    
            echo json_encode([
                'status' => 'ready'
            ]);
        }
    }

    Production health responses should avoid exposing credentials, internal addresses, stack traces, or sensitive dependency details.

    Scaling and Availability

    Horizontal scaling can improve availability only when instances are placed and operated to avoid shared failure points.

    Weak redundancy:
    
    Instance 1
    Instance 2
    Instance 3
    
    All on one physical host
    or one failure domain
    
    
    Stronger direction:
    
    Instances distributed across
    independent failure domains
    with healthy traffic routing

    Multiple instances are not sufficient when all instances depend on one unavailable database, network path, configuration service, or region.

    Scaling vs High Availability

    Goal Primary Question
    Scaling Can the system handle more workload?
    High availability Can the service continue after expected failures?
    Elasticity Can capacity follow changing demand efficiently?
    Disaster recovery Can service and data be restored after a major failure?

    Horizontal scaling can support high availability, but the two objectives must be designed and tested separately.

    Retry Amplification

    When the system becomes overloaded, aggressive retries can increase the workload.

    Service slows
          |
          v
    Requests time out
          |
          v
    Clients retry immediately
          |
          v
    Traffic increases
          |
          v
    Service slows further

    Safer retry behaviour includes:

    • Bounded attempts
    • Exponential backoff
    • Randomized jitter
    • Request deadlines
    • Idempotency
    • Retry guidance from the server where supported
    • Admission control or circuit breaking where appropriate

    Load Shedding

    A system can reject, delay, or degrade lower-priority work when available capacity is exhausted.

    Examples include:

    • Rejecting requests above a tenant quota
    • Serving an approved stale cached response
    • Deferring nonessential analytics
    • Reducing optional response detail
    • Queueing background work
    • Returning an explicit overload response

    Load shedding protects essential operations but does not replace capacity planning.

    Vertical and Horizontal Scaling Together

    Many production systems use both approaches.

    Initial stage:
    
    1 medium instance
    
    
    Growth stage:
    
    Scale up to a larger instance
    
    
    Availability stage:
    
    Run several instances
    
    
    Elastic stage:
    
    Automatically scale instance count
    
    
    Optimization stage:
    
    Adjust both instance size
    and instance count

    Using instances that are too small can create excessive coordination and overhead. Using instances that are too large can waste capacity and increase failure impact.

    Cost Considerations

    Scaling should be evaluated using total useful capacity rather than only instance price.

    A simplified compute-cost estimate is:

    \[ ComputeCost = InstanceCount \times CostPerInstance \times RunningDuration \]

    Also consider:

    • Load-balancer cost
    • Data-transfer cost
    • Storage and backup cost
    • Database scaling cost
    • Logging and monitoring cost
    • Minimum idle capacity
    • Autoscaling startup delay
    • Engineering and operational complexity

    Resource Utilization

    Scaling decisions should identify which resource is saturated.

    Symptom Possible Bottleneck
    High CPU and runnable work CPU-bound application logic
    High memory and swapping Memory pressure or leak
    Low CPU with slow database calls Database or network dependency
    High disk latency Storage I/O limit
    Long queue age Insufficient worker throughput or slow dependency
    Connection acquisition delay Connection-pool or database connection limit

    Bottleneck rule: Scaling the wrong resource wastes money. More application instances do not fix an inefficient query, exhausted database connection limit, memory leak, locked row, or unavailable dependency.

    Load Testing

    Load testing establishes the safe capacity and scaling behaviour of the system.

    A representative test should include:

    • Realistic request mix
    • Realistic database size
    • Authentication and authorization
    • Cache-hit and cache-miss paths
    • Expected concurrent users
    • Peak and burst traffic
    • Background jobs
    • Failure of selected instances
    • Scale-out and scale-in events
    • Downstream service limits
    • Sustained load after autoscaling

    A short test can measure initial capacity while missing memory growth, database saturation, queue accumulation, and autoscaling oscillation.

    Performance Measurements

    Measure:

    • Requests per second
    • Successful requests
    • Error and timeout rate
    • Median response time
    • High-percentile response time
    • CPU and memory by instance
    • Database latency and connections
    • Queue length and message age
    • Instance startup time
    • Scale-out completion time
    • Drain and scale-in time
    • Cost per successful workload unit

    Observability

    Useful scaling metrics include:

    • Current instance count
    • Desired instance count
    • Healthy backend count
    • Requests per instance
    • CPU and memory per instance
    • Request latency by instance
    • Error rate by instance
    • Load-balancer backend failures
    • Readiness and liveness failures
    • Autoscaling decisions
    • Instance provisioning failures
    • Database connection usage
    • Cache hit rate
    • Queue backlog
    • Scale-in termination duration
    • Cost by service and instance pool

    Alert Conditions

    Alert when:

    • Healthy instance count falls below the required minimum
    • Response time exceeds the service objective
    • Scaling reaches the configured maximum
    • New instances repeatedly fail readiness checks
    • Instance provisioning fails
    • Load is distributed unevenly
    • Database connections approach their safe limit
    • Queue backlog continues growing after scale-out
    • Scale-in repeatedly interrupts work
    • Autoscaling oscillates
    • One dependency remains saturated after application scaling

    Scaling Troubleshooting Workflow

    1. Capture the exact performance or availability symptom.
    2. Identify the affected component.
    3. Compare CPU, memory, disk, network, database, and queue metrics.
    4. Check whether the problem affects every instance or only selected instances.
    5. Check load-balancer distribution.
    6. Check readiness and health-check results.
    7. Check database connections and query latency.
    8. Check cache and session behaviour.
    9. Check queue backlog and worker throughput.
    10. Check autoscaling thresholds and cooldowns.
    11. Check instance startup and warm-up time.
    12. Check retry traffic and timeouts.
    13. Test the bottleneck under representative load.
    14. Scale the limiting component or correct the inefficient design.

    Common Scaling Mistakes

    1

    Scaling without Finding the Bottleneck

    More application capacity does not fix a slow database query, external dependency, disk limit, or lock contention.

    2

    Keeping Sessions in One Instance's Memory

    Requests fail when the load balancer routes the user to another instance.

    3

    Writing Shared Files to Local Disk

    Files handled by one instance are unavailable to the other instances.

    4

    Scaling Application Connections without a Database Budget

    Every new instance adds connection pools and can overwhelm the shared database.

    5

    Using CPU as the Only Scaling Signal

    Memory, queue delay, dependency latency, connections, or storage can be saturated while CPU remains moderate.

    6

    Scaling Out Too Late

    Instances become ready only after users are already experiencing high latency.

    7

    Scaling In Too Quickly

    Temporary traffic reductions cause capacity removal and repeated oscillation.

    8

    Terminating Instances without Draining

    In-flight requests, uploads, messages, and long-lived connections are interrupted.

    9

    Using Fragile Health Checks

    Expensive or overly dependent probes can remove healthy application instances during a downstream problem.

    10

    Assuming Multiple Instances Guarantee Availability

    All instances can still share one failure domain, database, network path, or configuration dependency.

    11

    Allowing Immediate Synchronized Retries

    Retry storms increase the workload during an existing overload condition.

    12

    Testing Only Steady Uniform Traffic

    The system remains untested against bursts, failures, startup delay, cache misses, deployments, and scale-in.

    Recommended Test Cases

    Test Expected Evidence
    Vertical scale-up The workload gains capacity and any restart impact is documented
    Horizontal scale-out New healthy instances receive traffic after readiness succeeds
    Load distribution Requests distribute according to the selected routing policy
    Backend failure The unhealthy instance is removed from active routing
    Readiness delay A starting instance receives no traffic before it is ready
    Scale-in Requests and background work drain before termination
    Session continuity Requests remain valid when routed to different instances
    Shared-file access Content uploaded through one instance is available through another
    Database connection budget Maximum instance count does not exhaust database connections
    Traffic burst Minimum capacity and scale-out delay remain within the objective
    Autoscaling oscillation Cooldown and separate thresholds prevent repeated scaling actions
    Queue backlog Worker scaling reduces oldest-message age without overwhelming dependencies
    Retry storm Backoff, jitter, and retry limits protect the service
    Maximum capacity reached Alerts and overload controls activate predictably

    Scaling Best Practices

    Recommended Practices

    • Measure the workload before selecting a scaling strategy.
    • Identify the limiting resource rather than scaling every component.
    • Use vertical scaling for simple, stateful, or non-distributable workloads where appropriate.
    • Use horizontal scaling when the workload can be distributed safely.
    • Design application instances to be replaceable.
    • Externalize sessions and shared files.
    • Protect shared databases with a connection budget.
    • Use load balancers and reliable readiness checks.
    • Keep health endpoints focused and inexpensive.
    • Maintain minimum capacity for failures and startup delay.
    • Use workload-relevant autoscaling signals.
    • Separate scale-out and scale-in thresholds.
    • Use cooldown periods to reduce oscillation.
    • Drain requests and messages before terminating instances.
    • Use bounded retries with exponential backoff and jitter.
    • Distribute instances across appropriate failure domains.
    • Test load balancing, scaling, and failure behaviour together.
    • Measure high-percentile latency, not only averages.
    • Monitor scaling limits and downstream dependencies.
    • Evaluate both performance benefit and total cost.

    Practice Exercise

    Design the scaling strategy for your online learning platform.

    Requirements

    1. Estimate normal and peak API request rates.
    2. Load-test one application instance.
    3. Identify CPU, memory, database, and network bottlenecks.
    4. Compare scaling up with scaling out.
    5. Place several instances behind a load balancer.
    6. Create liveness and readiness endpoints.
    7. Move sessions to a shared storage mechanism.
    8. Move uploaded course files to object storage.
    9. Set a database connection budget.
    10. Define minimum and maximum application capacity.
    11. Choose scale-out and scale-in signals.
    12. Add startup and warm-up handling.
    13. Add graceful request draining.
    14. Test one instance failure.
    15. Test a sudden traffic burst.
    16. Test the maximum-capacity condition.
    17. Measure performance and cost before and after scaling.

    Scaling-decision Template

    Component Scaling Direction Primary Signal Important Constraint
    Web/API tier Horizontal Request utilization and latency Statelessness and database connections
    Background workers Horizontal Queue length and oldest-message age Downstream capacity and idempotency
    Relational database Vertical first, then workload-specific read or data distribution CPU, storage latency, connections, and query performance Transactions, consistency, and recovery
    Cache Platform-specific horizontal or vertical scaling Memory, request rate, and eviction Key distribution and consistency
    Object storage Managed platform capacity Request and transfer metrics Lifecycle, authorization, and cost
    Search service Shard, replica, or node scaling according to platform Query latency, indexing rate, and storage Freshness and shard distribution

    Frequently Asked Questions

    1

    What is vertical scaling?

    Vertical scaling changes the capacity of an existing resource by increasing or reducing CPU, memory, storage, network, or service tier.

    2

    What is horizontal scaling?

    Horizontal scaling changes capacity by adding or removing resource instances.

    3

    What is scale up?

    Scale up means increasing the capacity of an existing instance.

    4

    What is scale out?

    Scale out means adding more instances to distribute workload.

    5

    Which approach is simpler?

    Vertical scaling is generally simpler because the application continues using fewer instances, but it has a maximum resource limit and can require interruption.

    6

    Why must horizontally scaled applications be stateless?

    Any healthy instance should be able to handle the next request without depending on private state stored only in another application process.

    7

    What is autoscaling?

    Autoscaling automatically adds or removes resources according to monitored conditions, schedules, and configured limits.

    8

    Why is a load balancer needed?

    A load balancer distributes requests among eligible backend instances and can remove unhealthy instances from active routing.

    9

    What is the difference between liveness and readiness?

    Liveness indicates whether the process can make progress. Readiness indicates whether the instance can safely receive new traffic.

    10

    Can the database become a bottleneck after scaling out?

    Yes. Additional application instances can create more connections, concurrent queries, locks, and storage traffic against the same database.

    11

    Does horizontal scaling guarantee high availability?

    No. Instances must be distributed across appropriate failure domains, and shared dependencies must also be resilient.

    12

    Can vertical and horizontal scaling be used together?

    Yes. Many systems select an appropriate instance size and then scale the number of those instances according to demand.

    Key Takeaway

    Vertical scaling increases the capacity of an existing instance and is often the simpler option for stateful, legacy, or non-distributable workloads. Horizontal scaling adds instances and can provide elasticity, redundancy, rolling deployment, and greater total capacity, but the application must support distributed execution. Externalize sessions and files, place healthy instances behind a load balancer, define focused readiness checks, budget database connections, and drain work safely during scale-in. Autoscaling should use workload-relevant signals and account for startup delay, warm-up, cooldowns, downstream limits, and failure headroom. Most importantly, identify the real bottleneck before scaling. Increasing the wrong resource raises cost without solving the performance problem.