L4 vs L7 balancing
Vertical vs Horizontal Scaling
Learn how vertical scaling increases the capacity of an existing server, how horizontal scaling adds more server instances, and how statelessness, load balancing, health checks, shared state, databases, queues, autoscaling, graceful shutdown, observability, and cost influence the correct scaling strategy.
Introduction
An application that performs well for a small number of users can slow down as traffic, stored data, background work, and concurrent requests increase.
The system can experience:
- Higher response time
- Increased CPU utilization
- Memory exhaustion
- Connection-pool saturation
- Longer queue processing time
- Database contention
- More timeouts and retries
- Reduced availability during failures
Scaling increases or reduces system capacity so the application can continue meeting its performance, reliability, and cost objectives under changing demand.
There are two primary scaling directions:
- Vertical scaling: Increase or decrease the capacity of an existing resource.
- Horizontal scaling: Add or remove resource instances.
Core idea: Vertical scaling makes one instance stronger. Horizontal scaling distributes work across several instances. Vertical scaling is usually simpler, while horizontal scaling can provide greater elasticity and resilience when the application is designed for distribution.
This lesson begins the Load Balancing, Proxies & Elastic Scaling module. The concepts introduced here provide the foundation for load balancers, reverse proxies, health checks, session management, redundancy, failover, and autoscaling.
Prerequisites
| # | Prerequisite | Why It Is Needed |
|---|---|---|
| 1 | CPU, memory, disk, and network fundamentals | Scaling decisions depend on the resource that limits the workload. |
| 2 | Processes and threads | Application concurrency affects how one instance uses available resources. |
| 3 | HTTP request lifecycle | Horizontal scaling distributes incoming requests across several backends. |
| 4 | Databases and connection pools | Adding application instances can increase pressure on shared databases. |
| 5 | Caching and sessions | Instance-local state can prevent requests from moving safely between servers. |
| 6 | Observability | Metrics are required to identify bottlenecks and validate scaling behaviour. |
What Is Scalability?
Scalability is the ability of a system to maintain acceptable behaviour when workload changes.
The workload can include:
- Concurrent users
- Requests per second
- Background jobs
- Messages waiting in queues
- Database queries
- Stored records
- Uploaded files
- Network traffic
A scalable system should have an understood way to increase capacity before a bottleneck causes unacceptable latency or failures.
Scalability vs Elasticity
| Concept | Meaning |
|---|---|
| Scalability | The system can handle additional workload by increasing resources. |
| Elasticity | The system can add and remove resources in response to changing demand. |
| Capacity | The amount of workload the current deployment can handle within its objectives. |
| Efficiency | The useful work produced from the allocated resources and cost. |
A system can be scalable without being automatically elastic. For example, operators can manually increase server capacity when growth occurs.
What Is Vertical Scaling?
Vertical scaling changes the capacity of an existing resource while keeping the number of instances conceptually unchanged.
Before vertical scaling:
1 application server
2 CPU cores
8 GB memory
After scaling up:
1 application server
8 CPU cores
32 GB memory
Increasing capacity is called scaling up. Reducing capacity is called scaling down.
Resources That Can Be Increased
- CPU cores
- Processor performance
- Memory
- Disk size
- Disk performance
- Network capacity
- Database compute tier
- Application-service plan size
Vertical-scaling Architecture
Clients
|
v
Application Server
Before:
+----------------------+
| 2 CPU |
| 8 GB RAM |
| Limited disk IOPS |
+----------------------+
Scale up
|
v
After:
+----------------------+
| 8 CPU |
| 32 GB RAM |
| Higher disk IOPS |
+----------------------+
The application architecture remains mostly unchanged. The existing instance receives more capacity.
Advantages of Vertical Scaling
- Usually simpler than distributing the application
- Requires fewer instances to configure and monitor
- Can work for stateful or legacy applications
- Can avoid distributed session and coordination concerns
- Can improve single-process or single-thread performance
- Can be an effective short-term response to gradual growth
- Can require fewer application-code changes
Limitations of Vertical Scaling
- Every resource type has an upper hardware or service-tier limit
- Scaling can require a restart, redeployment, or temporary unavailability
- One instance can remain a single failure point
- Larger instances can be expensive
- Unused capacity can remain allocated during low traffic
- One process might not use all added CPU cores effectively
- Scaling one component does not fix a bottleneck in another component
Vertical-scaling rule: Scaling up can provide immediate capacity, but it does not remove the maximum-instance limit or automatically improve availability.
What Is Horizontal Scaling?
Horizontal scaling changes capacity by adding or removing instances.
Before horizontal scaling:
1 application instance
After scaling out:
Application instance 1
Application instance 2
Application instance 3
Application instance 4
Adding instances is called scaling out. Removing instances is called scaling in.
Horizontal-scaling Architecture
+------------------+
Clients ------------>| Load Balancer |
+------------------+
| | |
+-------------+ | +-------------+
| | |
v v v
+-------------+ +-------------+ +-------------+
| App | | App | | App |
| Instance 1 | | Instance 2 | | Instance 3 |
+-------------+ +-------------+ +-------------+
A load balancer, gateway, proxy, service-discovery mechanism, or messaging system distributes work among the available instances.
Advantages of Horizontal Scaling
- Capacity can grow by adding instances
- Unhealthy instances can be removed from service
- Traffic can be distributed across several backends
- Instances can be added or removed elastically
- Rolling deployments become practical
- The system can tolerate selected instance failures
- Smaller commodity instances can replace one very large server
- Different regions or availability zones can host instances
Challenges of Horizontal Scaling
- The application must work correctly across several instances
- Instance-local sessions can cause routing problems
- Shared caches and databases can become bottlenecks
- Logs and traces are distributed
- Deployments require version compatibility
- Requests can arrive at different instances
- Background jobs can accidentally run more than once
- Coordination and distributed consistency become more complex
- Scale-in must not terminate instances with unfinished work
Vertical vs Horizontal Scaling
| Area | Vertical Scaling | Horizontal Scaling |
|---|---|---|
| Method | Increase the capacity of an existing instance | Add more instances |
| Common terms | Scale up and scale down | Scale out and scale in |
| Application changes | Often fewer changes | Can require statelessness and distributed coordination |
| Upper limit | Limited by the largest available resource size | Limited by architecture, coordination, shared dependencies, and service limits |
| Availability | One instance can remain a failure point | Several healthy instances can provide redundancy |
| Load distribution | Work remains on one larger instance | Work is distributed across instances |
| Scaling interruption | Can require restart or redeployment | New instances can often join while existing instances continue serving |
| Operational complexity | Usually lower | Usually higher |
| Elasticity | Less commonly automated | Commonly automated in cloud environments |
| Suitable workload | Stateful, legacy, or single-instance-oriented workload | Stateless or distributable workload |
Capacity Model
If one healthy application instance safely processes \(C\) requests per second and the expected peak workload is \(Q\), a simplified instance-count estimate is:
\[ RequiredInstances = \left\lceil \frac{Q}{C} \right\rceil \]
Additional capacity should be considered for failures, deployments, unexpected bursts, and measurement uncertainty.
Example
Tested safe capacity per instance:
500 requests per second
Expected peak traffic:
1,800 requests per second
Minimum calculated instances:
ceil(1800 / 500)
=
4 instances
This simplified estimate does not include failure headroom, uneven request cost, downstream limits, or high-percentile latency.
Failure Headroom
A horizontally scaled service should remain within acceptable performance after losing an expected number of instances.
Normal deployment:
4 instances
Failure scenario:
1 instance unavailable
Remaining capacity:
3 instances
Question:
Can 3 instances safely handle
the required workload?
Capacity planning should test the degraded scenario rather than assuming every instance is always available.
Load Balancer
A load balancer receives requests and distributes them across eligible backend instances.
Client request
|
v
Load balancer
|
+-- Backend 1
+-- Backend 2
+-- Backend 3
A load balancer can consider:
- Backend health
- Configured routing algorithm
- Connection count
- Backend capacity or weight
- Geographic location
- Protocol and request attributes
Exact features depend on the selected load balancer.
Basic Load-balancing Algorithms
| Algorithm | General Behaviour |
|---|---|
| Round robin | Distributes requests across backends in sequence |
| Weighted round robin | Sends more requests to backends assigned greater weight |
| Least connections | Prefers the backend with fewer active connections |
| Weighted least connections | Considers both active connections and backend weight |
| Hash-based routing | Maps a selected request value consistently to a backend |
| Resource-aware routing | Uses available backend health or utilization information |
The best algorithm depends on request duration, instance capacity, connection behaviour, and the availability of reliable backend metrics.
Health Checks
A load balancer should route traffic only to backends considered healthy.
Load balancer
|
v
Health-check request
|
+-- Healthy response:
| backend remains eligible
|
+-- Failed response:
backend is removed
from active rotation
Basic Health Endpoint
GET /health/live HTTP/1.1
Host: app.example.com
{
"status": "healthy"
}
A liveness check indicates whether the process is running. A readiness check indicates whether the instance is prepared to receive traffic.
Liveness vs Readiness
| Check | Question | Possible Action |
|---|---|---|
| Liveness | Is the process alive and capable of making progress? | Restart an unhealthy process |
| Readiness | Can the instance safely receive new requests? | Add or remove the instance from traffic rotation |
| Startup | Has initialization completed? | Delay other health decisions during startup |
Health-check rule: A process can be alive but not ready. Do not send production traffic to an instance until configuration, dependencies, migrations, and initialization required for request handling are complete.
Avoid Fragile Health Checks
A health endpoint should not become slow or unstable because it performs large queries or checks every downstream system.
Health request performs:
- Large database query
- External payment API call
- Search query
- Object-storage download
- Cache write
Liveness:
Process can make progress
Readiness:
Instance has required configuration
and can safely serve its core request path
Stateless Application Instances
Horizontal scaling is easier when any healthy instance can process the next request.
Request 1
-> Instance A
Request 2
-> Instance C
Request 3
-> Instance B
All requests remain correct.
This does not mean the complete application has no state. It means request-critical state is stored outside one specific application process or is included safely in the request.
The In-memory Session Problem
Login request
|
v
Instance A stores session in memory
Next request
|
v
Load balancer selects Instance B
Instance B:
Session not found
Possible solutions include:
- Store sessions in a shared key-value store
- Use appropriately protected self-contained tokens
- Use sticky sessions as a controlled transitional strategy
- Store authoritative session state in a shared database
Sticky Sessions
Sticky sessions attempt to route one client to the same backend instance.
User A
-> Instance 1
-> Instance 1
-> Instance 1
User B
-> Instance 2
-> Instance 2
-> Instance 2
Sticky routing can simplify migration from instance-local sessions, but it can cause uneven distribution and does not protect state when the selected instance fails.
Session rule: Sticky sessions can reduce routing flexibility. Externalizing important session state usually provides stronger horizontal-scaling and failover behaviour.
The Local-file Problem
Files written to one application instance might not be available on another instance.
Upload handled by Instance A
File stored at:
/var/app/uploads/lesson.pdf
Download handled by Instance B
Result:
File not found
Shared file storage, object storage, or another durable shared content system should hold files that must be accessible from multiple instances.
Shared Database Bottleneck
Adding application instances can shift the bottleneck to the database.
Before:
2 application instances
20 database connections each
=
40 possible connections
After scaling:
20 application instances
20 database connections each
=
400 possible connections
The application tier gained capacity, but the database can now face higher connection count, query concurrency, locking, CPU, and storage pressure.
Connection Budget
A simplified maximum connection estimate is:
\[ MaximumConnections = InstanceCount \times PoolSizePerInstance \]
Reserve capacity for administration, migrations, background jobs, and other database clients.
Connection-pool Management
Horizontal scaling should coordinate instance count and per-instance pool size.
Database connection budget:
200 connections
Maximum application instances:
10
Initial pool direction:
Less than 20 connections per instance
because capacity is also needed for:
- Background workers
- Administrative access
- Deployment operations
- Monitoring
- Failure headroom
Exact pool settings require workload measurement and database-specific guidance.
Multiple Scaling Layers
A production system contains several independently scalable components.
Client traffic
|
v
Gateway or load balancer
|
v
Web/API instances
|
v
Cache
|
v
Database
|
v
Object storage
Background flow:
Queue
|
v
Worker instances
Scaling the web tier does not automatically scale the cache, database, message broker, worker tier, or external dependencies.
Scaling Databases
Databases can be scaled vertically and horizontally, but horizontal database scaling is more complex than adding stateless application instances.
Vertical Database Scaling
- Increase CPU and memory
- Increase storage performance
- Increase service tier
- Increase connection or throughput capacity where supported
Horizontal Database Scaling
- Add read replicas
- Partition or shard data
- Separate workloads
- Create query-specific stores
- Use distributed database capabilities
Horizontal database scaling introduces replication, routing, consistency, partitioning, rebalancing, and cross-partition query considerations.
Read Replicas
+------------------+
Writes ------------>| Primary Database |
+------------------+
|
| Replication
v
+-------------+-------------+
| |
v v
+---------------+ +---------------+
| Read Replica 1| | Read Replica 2|
+---------------+ +---------------+
Read replicas can distribute suitable read traffic. Replication delay means a replica can temporarily return an older value.
Reads requiring immediate visibility after a write might need the authoritative primary or another supported consistency mechanism.
Queue-based Horizontal Scaling
Background processing can scale by adding consumers to a queue.
Producers
|
v
Message queue
|
+-- Worker 1
+-- Worker 2
+-- Worker 3
+-- Worker 4
The queue buffers work and consumers process available messages.
Worker scaling should consider:
- Queue length
- Oldest-message age
- Message processing time
- Downstream capacity
- Retry traffic
- Partition ordering
- Idempotency
Autoscaling
Autoscaling automatically changes resource capacity according to monitored conditions or schedules.
Metrics
|
v
Autoscaling decision
|
+-- Demand increased:
| add instances
|
+-- Demand decreased:
remove instances
Autoscaling commonly applies to horizontal compute scaling. Vertical resizing can be harder to automate because it can require restarting or redeploying a resource.
Autoscaling Signals
| Signal | Useful When | Limitation |
|---|---|---|
| CPU utilization | The workload is CPU-bound | Can miss memory, I/O, or dependency bottlenecks |
| Memory utilization | The workload consumes memory predictably | Can react poorly to leaks or cache behaviour |
| Request rate | Request cost is reasonably stable | Different requests can require different work |
| Response time | User latency reflects capacity pressure | Scaling can be too late after latency rises |
| Queue length | Workers process independent messages | Message duration can vary |
| Oldest-message age | Backlog delay is the important objective | Poison messages can distort the signal |
| Schedule | Demand follows a predictable calendar pattern | Unexpected traffic still requires another response |
Reactive vs Scheduled Scaling
Reactive Scaling
Load increases
|
v
Metric crosses threshold
|
v
Scaling action starts
|
v
New instance initializes
|
v
Health check succeeds
|
v
Instance receives traffic
Reactive scaling has a delay between detecting demand and receiving usable capacity.
Scheduled Scaling
Known examination period begins at 10:00
Scale out before 10:00
|
v
Instances become ready
|
v
Traffic arrives
Scheduled scaling is useful for predictable patterns and can be combined with metric-based scaling.
Instance Startup Time
Scaling decisions must account for how long a new instance requires before it becomes ready.
Instance provisioning
|
v
Runtime startup
|
v
Configuration loading
|
v
Dependency initialization
|
v
Application warm-up
|
v
Readiness success
|
v
Traffic begins
Slow startup can cause autoscaling to react too late. Maintain enough minimum capacity to absorb traffic while new instances initialize.
Warm-up Behaviour
A newly started instance can be healthy but not yet operating at full efficiency.
Warm-up can include:
- Loading application code
- Building runtime caches
- Opening database connections
- Loading configuration and secrets
- Compiling templates
- Initializing dependency clients
Gradual traffic ramp-up can protect a new instance from receiving too much work immediately.
Scaling Oscillation
Poorly tuned rules can repeatedly add and remove instances.
CPU rises
-> scale out
CPU falls
-> scale in
Traffic rises again
-> scale out
Traffic falls briefly
-> scale in
This behaviour is sometimes called oscillation or flapping.
Possible controls include:
- Separate scale-out and scale-in thresholds
- Cooldown periods
- Minimum instance lifetime
- Longer scale-in observation windows
- Minimum and maximum capacity
- Step scaling based on overload severity
Autoscaling rule: Scale out quickly enough to protect users, but scale in cautiously enough to avoid removing capacity during a temporary drop.
Graceful Scale-in
Removing an instance safely requires more than terminating the process.
Select instance for removal
|
v
Mark instance not ready
|
v
Stop assigning new requests
|
v
Allow in-flight work to finish
|
v
Stop background consumers
|
v
Close resources
|
v
Terminate instance
Long-running requests, uploads, WebSocket connections, and background jobs require explicit draining behaviour.
Long-lived Connections
WebSockets and other long-lived connections can complicate load balancing and scale-in.
Client
|
v
Long-lived connection
|
v
Instance A
Scale-in selects Instance A
|
v
Connection needs:
- Graceful closure
- Reconnection guidance
- State recovery
- Updated service discovery
Connection count can be a more useful scaling metric than simple request rate for long-lived protocols.
Conceptual Autoscaling Policy
autoscaling:
minimumInstances: 3
maximumInstances: 20
scaleOut:
metric: request-utilization
threshold: approved-high-threshold
evaluationWindow: approved-short-window
addInstances: 2
scaleIn:
metric: request-utilization
threshold: approved-low-threshold
evaluationWindow: approved-long-window
removeInstances: 1
cooldown:
afterScaleOut: approved-cooldown
afterScaleIn: approved-cooldown
readiness:
endpoint: /health/ready
termination:
drainRequests: true
gracePeriod: approved-drain-period
This is conceptual configuration. Thresholds and timing values must be established using representative workload tests and platform-specific guidance.
PHP Readiness Endpoint
<?php
declare(strict_types=1);
final class ReadinessController
{
public function __construct(
private ApplicationState $applicationState
) {
}
public function check(): void
{
header(
'Content-Type: application/json'
);
if (!$this->applicationState->isReady()) {
http_response_code(
503
);
echo json_encode([
'status' => 'not-ready'
]);
return;
}
http_response_code(
200
);
echo json_encode([
'status' => 'ready'
]);
}
}
Production health responses should avoid exposing credentials, internal addresses, stack traces, or sensitive dependency details.
Scaling and Availability
Horizontal scaling can improve availability only when instances are placed and operated to avoid shared failure points.
Weak redundancy:
Instance 1
Instance 2
Instance 3
All on one physical host
or one failure domain
Stronger direction:
Instances distributed across
independent failure domains
with healthy traffic routing
Multiple instances are not sufficient when all instances depend on one unavailable database, network path, configuration service, or region.
Scaling vs High Availability
| Goal | Primary Question |
|---|---|
| Scaling | Can the system handle more workload? |
| High availability | Can the service continue after expected failures? |
| Elasticity | Can capacity follow changing demand efficiently? |
| Disaster recovery | Can service and data be restored after a major failure? |
Horizontal scaling can support high availability, but the two objectives must be designed and tested separately.
Retry Amplification
When the system becomes overloaded, aggressive retries can increase the workload.
Service slows
|
v
Requests time out
|
v
Clients retry immediately
|
v
Traffic increases
|
v
Service slows further
Safer retry behaviour includes:
- Bounded attempts
- Exponential backoff
- Randomized jitter
- Request deadlines
- Idempotency
- Retry guidance from the server where supported
- Admission control or circuit breaking where appropriate
Load Shedding
A system can reject, delay, or degrade lower-priority work when available capacity is exhausted.
Examples include:
- Rejecting requests above a tenant quota
- Serving an approved stale cached response
- Deferring nonessential analytics
- Reducing optional response detail
- Queueing background work
- Returning an explicit overload response
Load shedding protects essential operations but does not replace capacity planning.
Vertical and Horizontal Scaling Together
Many production systems use both approaches.
Initial stage:
1 medium instance
Growth stage:
Scale up to a larger instance
Availability stage:
Run several instances
Elastic stage:
Automatically scale instance count
Optimization stage:
Adjust both instance size
and instance count
Using instances that are too small can create excessive coordination and overhead. Using instances that are too large can waste capacity and increase failure impact.
Cost Considerations
Scaling should be evaluated using total useful capacity rather than only instance price.
A simplified compute-cost estimate is:
\[ ComputeCost = InstanceCount \times CostPerInstance \times RunningDuration \]
Also consider:
- Load-balancer cost
- Data-transfer cost
- Storage and backup cost
- Database scaling cost
- Logging and monitoring cost
- Minimum idle capacity
- Autoscaling startup delay
- Engineering and operational complexity
Resource Utilization
Scaling decisions should identify which resource is saturated.
| Symptom | Possible Bottleneck |
|---|---|
| High CPU and runnable work | CPU-bound application logic |
| High memory and swapping | Memory pressure or leak |
| Low CPU with slow database calls | Database or network dependency |
| High disk latency | Storage I/O limit |
| Long queue age | Insufficient worker throughput or slow dependency |
| Connection acquisition delay | Connection-pool or database connection limit |
Bottleneck rule: Scaling the wrong resource wastes money. More application instances do not fix an inefficient query, exhausted database connection limit, memory leak, locked row, or unavailable dependency.
Load Testing
Load testing establishes the safe capacity and scaling behaviour of the system.
A representative test should include:
- Realistic request mix
- Realistic database size
- Authentication and authorization
- Cache-hit and cache-miss paths
- Expected concurrent users
- Peak and burst traffic
- Background jobs
- Failure of selected instances
- Scale-out and scale-in events
- Downstream service limits
- Sustained load after autoscaling
A short test can measure initial capacity while missing memory growth, database saturation, queue accumulation, and autoscaling oscillation.
Performance Measurements
Measure:
- Requests per second
- Successful requests
- Error and timeout rate
- Median response time
- High-percentile response time
- CPU and memory by instance
- Database latency and connections
- Queue length and message age
- Instance startup time
- Scale-out completion time
- Drain and scale-in time
- Cost per successful workload unit
Observability
Useful scaling metrics include:
- Current instance count
- Desired instance count
- Healthy backend count
- Requests per instance
- CPU and memory per instance
- Request latency by instance
- Error rate by instance
- Load-balancer backend failures
- Readiness and liveness failures
- Autoscaling decisions
- Instance provisioning failures
- Database connection usage
- Cache hit rate
- Queue backlog
- Scale-in termination duration
- Cost by service and instance pool
Alert Conditions
Alert when:
- Healthy instance count falls below the required minimum
- Response time exceeds the service objective
- Scaling reaches the configured maximum
- New instances repeatedly fail readiness checks
- Instance provisioning fails
- Load is distributed unevenly
- Database connections approach their safe limit
- Queue backlog continues growing after scale-out
- Scale-in repeatedly interrupts work
- Autoscaling oscillates
- One dependency remains saturated after application scaling
Scaling Troubleshooting Workflow
- Capture the exact performance or availability symptom.
- Identify the affected component.
- Compare CPU, memory, disk, network, database, and queue metrics.
- Check whether the problem affects every instance or only selected instances.
- Check load-balancer distribution.
- Check readiness and health-check results.
- Check database connections and query latency.
- Check cache and session behaviour.
- Check queue backlog and worker throughput.
- Check autoscaling thresholds and cooldowns.
- Check instance startup and warm-up time.
- Check retry traffic and timeouts.
- Test the bottleneck under representative load.
- Scale the limiting component or correct the inefficient design.
Common Scaling Mistakes
Scaling without Finding the Bottleneck
More application capacity does not fix a slow database query, external dependency, disk limit, or lock contention.
Keeping Sessions in One Instance's Memory
Requests fail when the load balancer routes the user to another instance.
Writing Shared Files to Local Disk
Files handled by one instance are unavailable to the other instances.
Scaling Application Connections without a Database Budget
Every new instance adds connection pools and can overwhelm the shared database.
Using CPU as the Only Scaling Signal
Memory, queue delay, dependency latency, connections, or storage can be saturated while CPU remains moderate.
Scaling Out Too Late
Instances become ready only after users are already experiencing high latency.
Scaling In Too Quickly
Temporary traffic reductions cause capacity removal and repeated oscillation.
Terminating Instances without Draining
In-flight requests, uploads, messages, and long-lived connections are interrupted.
Using Fragile Health Checks
Expensive or overly dependent probes can remove healthy application instances during a downstream problem.
Assuming Multiple Instances Guarantee Availability
All instances can still share one failure domain, database, network path, or configuration dependency.
Allowing Immediate Synchronized Retries
Retry storms increase the workload during an existing overload condition.
Testing Only Steady Uniform Traffic
The system remains untested against bursts, failures, startup delay, cache misses, deployments, and scale-in.
Recommended Test Cases
| Test | Expected Evidence |
|---|---|
| Vertical scale-up | The workload gains capacity and any restart impact is documented |
| Horizontal scale-out | New healthy instances receive traffic after readiness succeeds |
| Load distribution | Requests distribute according to the selected routing policy |
| Backend failure | The unhealthy instance is removed from active routing |
| Readiness delay | A starting instance receives no traffic before it is ready |
| Scale-in | Requests and background work drain before termination |
| Session continuity | Requests remain valid when routed to different instances |
| Shared-file access | Content uploaded through one instance is available through another |
| Database connection budget | Maximum instance count does not exhaust database connections |
| Traffic burst | Minimum capacity and scale-out delay remain within the objective |
| Autoscaling oscillation | Cooldown and separate thresholds prevent repeated scaling actions |
| Queue backlog | Worker scaling reduces oldest-message age without overwhelming dependencies |
| Retry storm | Backoff, jitter, and retry limits protect the service |
| Maximum capacity reached | Alerts and overload controls activate predictably |
Scaling Best Practices
Recommended Practices
- Measure the workload before selecting a scaling strategy.
- Identify the limiting resource rather than scaling every component.
- Use vertical scaling for simple, stateful, or non-distributable workloads where appropriate.
- Use horizontal scaling when the workload can be distributed safely.
- Design application instances to be replaceable.
- Externalize sessions and shared files.
- Protect shared databases with a connection budget.
- Use load balancers and reliable readiness checks.
- Keep health endpoints focused and inexpensive.
- Maintain minimum capacity for failures and startup delay.
- Use workload-relevant autoscaling signals.
- Separate scale-out and scale-in thresholds.
- Use cooldown periods to reduce oscillation.
- Drain requests and messages before terminating instances.
- Use bounded retries with exponential backoff and jitter.
- Distribute instances across appropriate failure domains.
- Test load balancing, scaling, and failure behaviour together.
- Measure high-percentile latency, not only averages.
- Monitor scaling limits and downstream dependencies.
- Evaluate both performance benefit and total cost.
Practice Exercise
Design the scaling strategy for your online learning platform.
Requirements
- Estimate normal and peak API request rates.
- Load-test one application instance.
- Identify CPU, memory, database, and network bottlenecks.
- Compare scaling up with scaling out.
- Place several instances behind a load balancer.
- Create liveness and readiness endpoints.
- Move sessions to a shared storage mechanism.
- Move uploaded course files to object storage.
- Set a database connection budget.
- Define minimum and maximum application capacity.
- Choose scale-out and scale-in signals.
- Add startup and warm-up handling.
- Add graceful request draining.
- Test one instance failure.
- Test a sudden traffic burst.
- Test the maximum-capacity condition.
- Measure performance and cost before and after scaling.
Scaling-decision Template
| Component | Scaling Direction | Primary Signal | Important Constraint |
|---|---|---|---|
| Web/API tier | Horizontal | Request utilization and latency | Statelessness and database connections |
| Background workers | Horizontal | Queue length and oldest-message age | Downstream capacity and idempotency |
| Relational database | Vertical first, then workload-specific read or data distribution | CPU, storage latency, connections, and query performance | Transactions, consistency, and recovery |
| Cache | Platform-specific horizontal or vertical scaling | Memory, request rate, and eviction | Key distribution and consistency |
| Object storage | Managed platform capacity | Request and transfer metrics | Lifecycle, authorization, and cost |
| Search service | Shard, replica, or node scaling according to platform | Query latency, indexing rate, and storage | Freshness and shard distribution |
Frequently Asked Questions
What is vertical scaling?
Vertical scaling changes the capacity of an existing resource by increasing or reducing CPU, memory, storage, network, or service tier.
What is horizontal scaling?
Horizontal scaling changes capacity by adding or removing resource instances.
What is scale up?
Scale up means increasing the capacity of an existing instance.
What is scale out?
Scale out means adding more instances to distribute workload.
Which approach is simpler?
Vertical scaling is generally simpler because the application continues using fewer instances, but it has a maximum resource limit and can require interruption.
Why must horizontally scaled applications be stateless?
Any healthy instance should be able to handle the next request without depending on private state stored only in another application process.
What is autoscaling?
Autoscaling automatically adds or removes resources according to monitored conditions, schedules, and configured limits.
Why is a load balancer needed?
A load balancer distributes requests among eligible backend instances and can remove unhealthy instances from active routing.
What is the difference between liveness and readiness?
Liveness indicates whether the process can make progress. Readiness indicates whether the instance can safely receive new traffic.
Can the database become a bottleneck after scaling out?
Yes. Additional application instances can create more connections, concurrent queries, locks, and storage traffic against the same database.
Does horizontal scaling guarantee high availability?
No. Instances must be distributed across appropriate failure domains, and shared dependencies must also be resilient.
Can vertical and horizontal scaling be used together?
Yes. Many systems select an appropriate instance size and then scale the number of those instances according to demand.
Key Takeaway
Vertical scaling increases the capacity of an existing instance and is often the simpler option for stateful, legacy, or non-distributable workloads. Horizontal scaling adds instances and can provide elasticity, redundancy, rolling deployment, and greater total capacity, but the application must support distributed execution. Externalize sessions and files, place healthy instances behind a load balancer, define focused readiness checks, budget database connections, and drain work safely during scale-in. Autoscaling should use workload-relevant signals and account for startup delay, warm-up, cooldowns, downstream limits, and failure headroom. Most importantly, identify the real bottleneck before scaling. Increasing the wrong resource raises cost without solving the performance problem.