autoscaling
Autoscaling
Learn how autoscaling observes workload signals, calculates required capacity, adds or removes application instances, integrates with load balancers and health checks, handles startup and shutdown safely, prevents scaling oscillation, protects downstream services, and balances performance, availability, and cost.
Introduction
Application traffic is rarely constant. An online learning platform can receive ordinary traffic during most of the day and substantially higher traffic during course launches, examinations, live sessions, campaigns, or scheduled learning activities.
If the deployment always runs with minimum capacity, sudden demand can cause:
- High response time
- Request timeouts
- Queue backlog
- CPU or memory saturation
- Connection-pool exhaustion
- Reduced application availability
If the deployment always runs with maximum capacity, much of that capacity can remain unused during low-demand periods.
Autoscaling automatically changes allocated capacity according to observed metrics, schedules, forecasts, or configured policies.
Core idea: Autoscaling is a feedback-control process. It observes demand, compares that demand with a target, calculates desired capacity, applies a scaling action, and then waits for the system to stabilize before making further decisions.
Observe workload
|
v
Compare metric with target
|
v
Calculate desired capacity
|
+-- Demand increased:
| scale out
|
+-- Demand decreased:
scale in
|
v
Wait for new capacity to stabilize
|
v
Observe again
Prerequisites
| # | Prerequisite | Why It Is Needed |
|---|---|---|
| 1 | Vertical and horizontal scaling | Autoscaling commonly changes the number of horizontally scaled instances. |
| 2 | Load balancing | New healthy instances must begin receiving traffic, while removed instances must leave routing safely. |
| 3 | Health checks | An instance should not receive production traffic before it becomes ready. |
| 4 | Statelessness and sessions | Instances must be replaceable without losing required user state. |
| 5 | Metrics and observability | Autoscaling decisions depend on accurate workload and performance signals. |
| 6 | Queues and background workers | Worker capacity is commonly scaled from backlog and processing-delay signals. |
| 7 | Databases and connection pools | Adding application instances can increase pressure on shared dependencies. |
What Is Autoscaling?
Autoscaling is the automatic adjustment of computing capacity to match changing workload requirements.
When demand increases, the platform can add resources to maintain the required performance. When demand decreases, unnecessary resources can be removed to reduce cost.
Low demand:
3 application instances
Traffic increases:
3 -> 5 -> 8 instances
Traffic decreases:
8 -> 6 -> 4 -> 3 instances
Autoscaling is most commonly associated with horizontal scaling, where the system adds or removes instances. Vertical resource changes can also be automated in some environments, but resizing can require restart or redeployment and is therefore usually less dynamic.
Scaling Terms
| Term | Meaning |
|---|---|
| Scale out | Add more instances |
| Scale in | Remove instances |
| Scale up | Increase the resources assigned to an existing instance |
| Scale down | Reduce the resources assigned to an existing instance |
| Desired capacity | The number of instances the scaling controller currently wants |
| Minimum capacity | The lowest permitted number of instances |
| Maximum capacity | The highest permitted number of instances |
| Cooldown | A stabilization period after a scaling action |
Scalability vs Elasticity vs Autoscaling
| Concept | Meaning |
|---|---|
| Scalability | The system can handle more load by receiving additional capacity. |
| Elasticity | Capacity can grow and shrink as demand changes. |
| Autoscaling | Policies and controllers adjust capacity automatically. |
An application can be horizontally scalable but not automatically scaled. Operators can still add and remove instances manually.
Components of an Autoscaling System
A complete autoscaling strategy contains several parts:
- Instrumentation that captures workload and resource metrics
- A monitoring system that collects and aggregates those metrics
- A scaling policy containing limits, targets, and rules
- A controller that calculates desired capacity
- A provisioning system that creates or removes resources
- Health checks that determine when new resources are usable
- A load balancer that updates backend membership
- Observability that records scaling decisions and outcomes
Application and infrastructure metrics
|
v
Monitoring system
|
v
Autoscaling policy
|
v
Scaling controller
|
v
Provision or remove instances
|
v
Health and readiness checks
|
v
Load-balancer pool
Scaling Signals
A scaling signal is a measured value used to decide whether capacity should change.
| Signal | Suitable Workload | Main Limitation |
|---|---|---|
| CPU utilization | CPU-bound request or computation processing | Does not directly identify memory, I/O, or dependency bottlenecks |
| Memory utilization | Workloads whose memory use follows demand predictably | Caches and memory leaks can distort the signal |
| Requests per second | Request costs are relatively predictable | Different endpoints can require substantially different work |
| Concurrent requests | Operations remain active for meaningful periods | One expensive request can differ from many inexpensive requests |
| Response time | User latency reflects capacity pressure | Scaling begins after degradation has already appeared |
| Queue length | Workers process queued background jobs | Message processing times can vary |
| Oldest-message age | Backlog waiting time is the important objective | A blocked or poison message can distort the measurement |
| Active connections | WebSocket, streaming, database, or long-lived TCP workloads | Connection activity and resource cost can vary |
| Schedule | Traffic follows a known recurring pattern | Unexpected demand still requires reactive protection |
CPU-based Autoscaling
CPU utilization is a common autoscaling signal because it is widely available and simple to observe.
Average CPU remains above
the approved upper target
|
v
Add application instances
Average CPU remains below
the approved lower target
|
v
Remove application instances cautiously
CPU-based scaling is useful only when adding instances distributes the CPU-producing work.
Additional application instances do not correct:
- A slow database query
- A shared database lock
- An exhausted downstream rate limit
- A memory leak
- A hot partition
- A blocked external service
Memory-based Autoscaling
Memory can be useful when demand produces predictable per-request or per-worker memory consumption.
Memory requires careful interpretation because:
- Runtime memory can remain allocated after traffic decreases
- Local caches can intentionally consume available memory
- A leak can cause continuous scale-out without recovery
- New instances can repeat the same memory problem
Metric rule: Select a signal that changes because of useful workload demand. Autoscaling should not hide a leak, inefficient query, dependency failure, or defective application.
Queue-based Autoscaling
Background workers can scale according to queue backlog.
Producers
|
v
Durable queue
|
+-- Worker 1
+-- Worker 2
+-- Worker 3
When backlog grows, more workers can be added. When the queue remains empty or within the approved lower range, workers can be removed.
Queue length alone can be misleading. A queue containing 1,000 one-second jobs differs greatly from a queue containing 1,000 ten-minute jobs.
Backlog-per-worker
A simplified signal is:
\[ BacklogPerWorker = \frac{ VisibleMessages }{ ActiveWorkers } \]
A time-oriented estimate can also consider average processing duration:
\[ EstimatedDrainTime = \frac{ QueueLength \times AverageProcessingTime }{ ActiveWorkers } \]
These formulas are planning approximations. Retries, message variation, ordering, dependency limits, and concurrency affect actual performance.
Main Autoscaling Approaches
| Approach | How It Works | Best Fit |
|---|---|---|
| Threshold-based | Adds or removes capacity when a metric crosses configured limits | Simple and predictable workloads |
| Target tracking | Adjusts capacity to keep a metric near a target value | Workload metrics that correlate with required capacity |
| Step scaling | Changes capacity by different amounts according to overload severity | Workloads where small and severe overloads need different responses |
| Scheduled scaling | Changes capacity at configured times | Predictable events or recurring traffic patterns |
| Predictive scaling | Uses observed history or forecasts to provision capacity before expected demand | Workloads with sufficiently repeatable demand patterns |
Threshold-based Scaling
If CPU remains above
the upper threshold:
Add capacity
If CPU remains below
the lower threshold:
Remove capacity
Use separate upper and lower thresholds. This creates a stable range and reduces repeated scale-out and scale-in actions around one boundary.
Scale-out threshold:
Higher utilization boundary
Stable operating range:
No scaling action
Scale-in threshold:
Lower utilization boundary
Target-tracking Scaling
Target tracking attempts to keep a selected metric near a configured target.
A simplified desired-capacity estimate is:
\[ DesiredInstances = CurrentInstances \times \frac{ CurrentMetric }{ TargetMetric } \]
Example
Current instances:
4
Current average utilization:
80%
Target utilization:
50%
Estimated desired instances:
4 × 80 / 50
=
6.4
Controller applies platform-specific
rounding and policy limits.
Exact calculations, rounding, stabilization, and evaluation behaviour depend on the autoscaling platform.
Step Scaling
Small overload:
Add 1 instance
Moderate overload:
Add 3 instances
Severe overload:
Add 6 instances
Step scaling can react more strongly to severe demand but must remain within the configured maximum capacity and downstream limits.
Scheduled Scaling
Scheduled scaling changes capacity before or during a predictable event.
Known examination begins:
10:00
Scheduled scale-out begins:
Before traffic arrives
At 10:00:
Additional instances are already
healthy and receiving traffic
Scheduled scaling is useful because reactive scaling requires time to detect load, provision resources, initialize the application, and pass readiness checks.
Predictive Scaling
Predictive scaling attempts to provision capacity before expected workload based on historical or forecast information.
Historical usage pattern
|
v
Forecast future demand
|
v
Provision capacity in advance
|
v
Validate actual demand
|
v
Reactive rules handle differences
Forecasts can be wrong. Predictive scaling should operate within minimum, maximum, cost, and safety boundaries and should be complemented by reactive protection.
Startup Delay
A scaling action does not create usable capacity immediately.
Scaling decision
|
v
Provision infrastructure
|
v
Start runtime
|
v
Load application
|
v
Load configuration and secrets
|
v
Initialize dependencies
|
v
Warm required components
|
v
Pass readiness check
|
v
Join load-balancer pool
The full period from scaling decision to usable capacity is the effective scale-out delay.
Capacity rule: Autoscaling cannot replace minimum headroom. Existing capacity must continue serving users while additional instances are starting.
Application Warm-up
A newly started instance can pass a basic process check before reaching normal production efficiency.
Warm-up can include:
- Creating database connections
- Initializing runtime components
- Loading application configuration
- Loading templates or metadata
- Building safe local caches
- Establishing downstream clients
Readiness should prevent premature traffic. Gradual traffic ramp-up can also protect a new instance where supported.
Cooldown and Stabilization
A cooldown or stabilization period gives a previous scaling action time to affect the observed workload.
Scale out
|
v
New instances starting
|
v
Metric remains temporarily high
|
v
Without stabilization:
Controller scales out repeatedly
With stabilization:
Controller waits for the previous
capacity change to take effect
Cooldown should account for provisioning, startup, readiness, connection establishment, cache warm-up, and metric-collection delay.
Scaling Oscillation
Oscillation, also called flapping, occurs when capacity repeatedly scales out and in.
Load rises
-> Scale out
Load falls briefly
-> Scale in
Capacity falls
-> Load rises
Scale out again
Possible controls include:
- Separate scale-out and scale-in thresholds
- Longer scale-in observation periods
- Cooldown or stabilization windows
- Minimum instance lifetimes
- Smaller scale-in steps
- Maximum scale-in rate
- Workload-aware metrics
Scaling-direction rule: Scale out quickly enough to protect performance, but scale in conservatively enough to avoid removing capacity during a temporary reduction in demand.
Minimum Capacity
Minimum capacity is the lowest instance count the autoscaling system can maintain.
It should consider:
- Ordinary traffic
- Expected instance failure
- Deployment activity
- Startup delay
- Availability-zone or failure-domain requirements
- Maintenance operations
- Sudden traffic before scaling completes
Minimum capacity should not be selected only from the lowest observed traffic.
Maximum Capacity
Maximum capacity prevents uncontrolled infrastructure growth.
It should consider:
- Cost boundaries
- Database connection capacity
- Downstream API limits
- Queue and broker capacity
- Network and load-balancer limits
- Available address space
- Service quotas
- Licensing constraints
Reaching maximum capacity should generate a clear alert because additional demand can no longer be addressed by scaling the configured resource group.
Downstream Bottlenecks
Scaling one tier increases demand on its dependencies.
Before scaling:
4 API instances
20 database connections each
=
80 potential connections
After scaling:
20 API instances
20 database connections each
=
400 potential connections
The API tier gained capacity, but the database can now face connection, query, CPU, memory, lock, and storage pressure.
Connection Budget
A simplified connection estimate is:
\[ MaximumConnections = MaximumInstances \times PoolSizePerInstance \]
Reserve capacity for workers, reporting, administration, deployment, monitoring, and recovery activity.
Hot Partitions and Autoscaling
Autoscaling does not automatically fix uneven workload distribution.
10 application instances
|
v
All requests target
one database partition
|
v
Hot partition remains
the system bottleneck
Add instances only when the workload can be distributed. Hot keys, serialized operations, global counters, and locked resources can remain bottlenecks after compute scale-out.
Safe Scale-in
Removing capacity requires graceful shutdown.
Select instance for scale-in
|
v
Mark readiness as false
|
v
Load balancer stops new requests
|
v
Drain in-flight requests
|
v
Stop message consumption
|
v
Complete or release owned work
|
v
Close connections
|
v
Terminate instance
Scale-in policy must consider:
- Long-running requests
- Uploads
- WebSocket connections
- Background jobs
- Queue visibility timeouts
- Local temporary processing
- Connection-drain duration
Autoscaling and Sessions
Required sessions should not exist only in one instance's memory.
Scale-in removes Instance A
|
v
User's next request reaches Instance C
|
v
Instance C retrieves the session
from a shared store
|
v
User workflow continues
Sticky sessions complicate scale-in because active users can remain bound to an instance selected for removal.
Autoscaling and Health Checks
New instances should enter the load-balancer pool only after readiness succeeds.
New instance created
|
v
Startup check
|
v
Application initialization
|
v
Readiness succeeds
|
v
Load balancer adds backend
An instance that cannot initialize should not count as usable capacity merely because it exists.
Failure Headroom
Capacity planning should account for expected failures.
Required normal capacity:
6 instances
One instance unavailable:
5 healthy instances remain
Question:
Can 5 instances safely handle
the required workload while a
replacement is starting?
Running every instance close to saturation leaves little capacity for failure, deployment, or sudden traffic.
Load Shedding
Autoscaling is not instantaneous. During severe demand, the system can still require overload protection.
Possible controls include:
- Rate limiting
- Concurrency limiting
- Bounded queues
- Admission control
- Serving approved stale cached data
- Deferring nonessential work
- Reducing optional response details
- Returning explicit overload responses
Do not silently degrade correctness-critical operations such as payment, inventory reservation, authorization changes, or financial posting.
Retry Amplification
Application becomes slow
|
v
Clients time out
|
v
Clients retry
|
v
Traffic increases
|
v
Autoscaling detects additional load
|
v
New capacity starts too late
or downstream system becomes overloaded
Use:
- Bounded retry attempts
- Exponential backoff
- Randomized jitter
- End-to-end deadlines
- Idempotency
- Overload-aware admission control
Web-tier Example
Client traffic
|
v
Application Load Balancer
|
+-- Web Instance 1
+-- Web Instance 2
+-- Web Instance 3
|
v
Shared services
Autoscaling signal:
Requests per healthy instance
Scale out:
Add web instances
Scale in:
Drain and remove web instances
The application instances should be stateless, health checked, and protected by minimum and maximum capacity.
Worker-tier Example
Course-video uploads
|
v
Processing queue
|
+-- Worker 1
+-- Worker 2
+-- Worker 3
Scaling signals:
- Visible messages
- Oldest-message age
- Processing duration
- Downstream capacity
Worker scale-out should not exceed the safe concurrency of object storage, media-processing dependencies, databases, or third-party APIs.
Conceptual Autoscaling Policy
autoscaling:
minimumInstances: 3
maximumInstances: 20
defaultInstances: 4
metric:
name: requests-per-healthy-instance
target: approved-target
scaleOut:
evaluationWindow: approved-short-window
maximumStep: approved-scale-out-step
scaleIn:
evaluationWindow: approved-long-window
maximumStep: approved-scale-in-step
stabilization:
afterScaleOut: approved-scale-out-cooldown
afterScaleIn: approved-scale-in-cooldown
readiness:
path: /health/ready
termination:
markNotReady: true
drainRequests: true
gracePeriod: approved-drain-period
This is conceptual configuration. Metric names, thresholds, timing, and policy structures depend on the selected platform and tested workload.
Kubernetes Horizontal Pod Autoscaling
Kubernetes can use a Horizontal Pod Autoscaler to adjust the desired replica count of a supported workload according to observed resource or custom metrics.
Conceptual HPA Definition
apiVersion: autoscaling/v2
kind: HorizontalPodAutoscaler
metadata:
name: course-api
spec:
scaleTargetRef:
apiVersion: apps/v1
kind: Deployment
name: course-api
minReplicas: 3
maxReplicas: 20
metrics:
- type: Resource
resource:
name: cpu
target:
type: Utilization
averageUtilization: approved-target
Resource requests, metrics availability, readiness, application startup, and node capacity must also be configured correctly. Adding pods does not help when there is no cluster capacity on which to schedule them.
Pod Scaling vs Node Scaling
| Layer | Scaling Action | Important Dependency |
|---|---|---|
| Application or pod layer | Add or remove application replicas | Enough compute capacity must already exist |
| Node or cluster layer | Add or remove worker machines | Infrastructure provisioning and scheduling |
| Database layer | Resize, replicate, or partition according to platform | Consistency, routing, storage, and recovery |
These layers can have different startup times and policies. Application replicas can remain pending while new compute nodes are being created.
Database Autoscaling
Database scaling differs from adding stateless application instances.
Possible database scaling actions include:
- Increase compute or memory
- Add read replicas
- Increase provisioned throughput
- Expand storage
- Rebalance or partition data
Review transaction behaviour, replication lag, connection routing, consistency, scale-in safety, and recovery before automating a database capacity change.
Cost Considerations
Autoscaling can reduce unused capacity, but it does not guarantee minimum cost.
A simplified compute-cost model is:
\[ ComputeCost = \sum \left( InstanceCount \times InstancePrice \times RunningDuration \right) \]
Also consider:
- Minimum always-on capacity
- Load-balancer cost
- Metrics and logging cost
- Data-transfer cost
- Database and cache scaling
- Startup and warm-up inefficiency
- Rapid scaling oscillation
- Reserved or committed capacity
- Licensing
Security Considerations
Automatically created instances must receive the same approved security configuration as existing instances.
Ensure:
- Instances are created from approved images or artifacts
- Secrets are retrieved through approved mechanisms
- Network restrictions apply automatically
- Identity and permissions use least privilege
- Monitoring and security tooling initialize correctly
- Temporary instances do not retain sensitive local state
- Scale-in follows secure cleanup and lifecycle policies
Observability
Useful autoscaling metrics include:
- Current instance count
- Desired instance count
- Healthy and ready instance count
- Scaling actions by direction
- Scaling-action failures
- Instance provisioning time
- Application startup time
- Readiness delay
- CPU and memory per instance
- Requests per instance
- Queue backlog per worker
- Oldest-message age
- Connection-pool usage
- Scale-in drain duration
- Time spent at maximum capacity
- Cost by service and scaling group
Scaling-event Record
{
"service": "course-api",
"action": "scale-out",
"previousCapacity": 4,
"desiredCapacity": 6,
"metric": "requests-per-healthy-instance",
"reason": "target exceeded",
"policyVersion": "approved-policy-version"
}
Scaling records should help operators understand why capacity changed without including credentials or sensitive request data.
Alert Conditions
Alert when:
- The service reaches maximum capacity
- Desired capacity cannot be provisioned
- New instances repeatedly fail readiness
- Scale-out does not reduce latency or backlog
- Autoscaling oscillates repeatedly
- Healthy capacity falls below the required minimum
- Database connections approach their safe limit
- Queue backlog grows after worker scale-out
- Scale-in interrupts active work
- One failure domain loses too much capacity
- Scaling cost grows unexpectedly
- Metrics required by the scaling controller are unavailable
Troubleshooting Workflow
- Identify the user-visible performance or availability symptom.
- Confirm the current and desired capacity.
- Identify which scaling policy is active.
- Inspect the metric that triggered or failed to trigger scaling.
- Check metric delay, aggregation, and missing data.
- Check minimum, maximum, and default capacity.
- Check cooldown and stabilization behaviour.
- Check provisioning failures and service quotas.
- Check application startup and readiness.
- Check load-balancer backend membership.
- Check database, cache, queue, and external dependency capacity.
- Check retry amplification and overload controls.
- Check graceful scale-in and connection draining.
- Compare the result with representative load-test evidence.
Common Autoscaling Mistakes
Selecting the Wrong Metric
Capacity changes without correcting the actual user-facing or processing bottleneck.
Using CPU as the Only Signal
Database latency, queue age, memory, connections, or storage can be saturated while CPU remains moderate.
Scaling Out Too Late
Traffic exceeds capacity before new instances finish provisioning and startup.
Scaling In Too Aggressively
Capacity is removed during a temporary reduction and must be recreated immediately afterward.
Using the Same Boundary for Scale-out and Scale-in
Minor metric variation can cause repeated scaling oscillation.
Ignoring Startup and Warm-up Time
Newly created resources exist but do not provide usable capacity soon enough.
Counting Unready Instances as Capacity
The scaling controller believes adequate capacity exists while the load balancer has too few usable backends.
Ignoring Downstream Limits
Additional instances overwhelm database connections, queues, caches, or external APIs.
Using Autoscaling to Hide a Memory Leak
The application repeatedly adds instances without correcting the defect.
Removing Instances without Draining
Scale-in interrupts requests, uploads, messages, and long-lived connections.
Setting Maximum Capacity without an Alert
The service stops scaling while operators remain unaware that no additional capacity is available.
Testing Only Uniform Traffic
The policy remains untested against bursts, hot keys, expensive routes, retries, cache misses, and dependency failures.
Recommended Test Cases
| Test | Expected Evidence |
|---|---|
| Normal scale-out | New instances become ready and receive traffic |
| Normal scale-in | Instances drain before termination |
| Sudden traffic burst | Minimum headroom protects users while new capacity starts |
| Sustained traffic increase | Capacity converges to the required level |
| Temporary metric spike | The policy avoids unnecessary repeated scaling |
| Traffic reduction | Scale-in occurs only after the approved stabilization period |
| Maximum capacity | The platform raises an alert and applies overload protection |
| Instance startup failure | The failed instance receives no production traffic |
| Readiness delay | New capacity is counted only after becoming usable |
| One instance failure | Remaining capacity serves traffic while replacement occurs |
| Database connection limit | Maximum application scale does not exhaust the database budget |
| Queue backlog | Worker scale-out reduces processing delay safely |
| Poison message | One unprocessable message does not trigger uncontrolled scaling |
| Retry storm | Backoff, jitter, and admission control protect capacity |
| Long-lived connections | Scale-in drains or reconnects clients according to policy |
| Autoscaling oscillation | Threshold separation and stabilization prevent flapping |
Autoscaling Best Practices
Recommended Practices
- Confirm that the workload can scale horizontally.
- Keep scalable application instances stateless and replaceable.
- Select metrics that correlate with useful workload demand.
- Use queue delay or backlog for asynchronous workers.
- Maintain meaningful minimum and maximum capacity.
- Retain headroom for startup delay and expected failures.
- Use separate scale-out and scale-in conditions.
- Scale out faster than scale in when availability is the priority.
- Use cooldown and stabilization windows.
- Count only healthy and ready instances as usable capacity.
- Warm new instances before sending significant traffic.
- Budget database and downstream connections at maximum scale.
- Limit worker concurrency according to downstream capacity.
- Use scheduled scaling for known demand where appropriate.
- Keep reactive rules for unexpected demand.
- Drain requests, jobs, and connections before scale-in.
- Use rate limiting and load shedding while capacity starts.
- Alert when maximum capacity is reached.
- Record every scaling decision and policy version.
- Load-test scale-out, scale-in, failures, and downstream saturation.
Practice Exercise
Design autoscaling for your online learning platform's API and video-processing workers.
Requirements
- Measure the safe capacity of one API instance.
- Estimate normal and peak request rates.
- Select minimum and maximum API capacity.
- Choose a workload metric for API scale-out.
- Define separate scale-out and scale-in conditions.
- Measure application startup and readiness time.
- Maintain capacity while new instances initialize.
- Set a maximum database-connection budget.
- Create a worker autoscaling policy based on queue backlog and age.
- Limit workers according to media-processing dependency capacity.
- Add graceful request and message draining.
- Add rate limiting for severe API demand.
- Test a sudden examination-period traffic burst.
- Test a poison message in the processing queue.
- Test one failed instance during peak load.
- Test the maximum-capacity scenario.
- Measure performance, failure behaviour, and cost.
Autoscaling-design Template
| Workload | Scaling Signal | Capacity Boundary | Primary Safety Control |
|---|---|---|---|
| Web and API requests | Requests per healthy instance and latency | Minimum and maximum ready instances | Readiness, connection budget, and rate limiting |
| Video-processing workers | Queue backlog and oldest-message age | Maximum safe processing concurrency | Durable queue, idempotency, and dependency protection |
| WebSocket service | Active connections and connection growth | Connection capacity per instance | Reconnection and graceful draining |
| Database readers | Read load, connection use, and replica capacity | Platform and consistency limits | Replication-lag and routing controls |
| Scheduled examination traffic | Known event schedule plus reactive metrics | Forecast capacity and cost boundary | Pre-scaling and maximum-capacity alerts |
Frequently Asked Questions
What is autoscaling?
Autoscaling automatically adjusts resource capacity according to observed metrics, schedules, forecasts, and configured limits.
What is scale out?
Scale out means adding more instances to distribute workload.
What is scale in?
Scale in means removing instances when less capacity is required.
Which metric should autoscaling use?
Use a metric that reliably reflects the workload and capacity pressure of the specific component, such as CPU, requests per instance, queue age, backlog, or active connections.
Why is minimum capacity required?
Minimum capacity serves ordinary traffic and provides headroom while new resources are provisioned or existing instances fail.
Why is maximum capacity required?
Maximum capacity limits cost and protects databases, external services, network resources, and other dependencies from uncontrolled concurrency.
What is a cooldown period?
A cooldown gives a previous scaling action time to affect the workload before another scaling decision is made.
What is scaling oscillation?
Scaling oscillation is repeated scale-out and scale-in caused by unstable thresholds, delayed metrics, or insufficient stabilization.
Can autoscaling fix a slow database query?
No. Additional application instances can increase database pressure without correcting the inefficient query.
Should autoscaling use CPU only?
Not necessarily. CPU is useful for CPU-bound workloads, but user latency, requests, memory, queue delay, connections, or custom metrics can better represent other workloads.
How should instances be removed?
Mark the instance not ready, stop new work, drain existing requests and messages, close connections, and then terminate it.
Does autoscaling guarantee availability?
No. The system still requires failure headroom, healthy routing, resilient dependencies, overload protection, tested startup, and safe scale-in.
Key Takeaway
Autoscaling automatically adjusts capacity to match changing demand. A reliable autoscaling design begins with a horizontally scalable workload and uses metrics that reflect actual capacity pressure. It defines minimum, maximum, and desired capacity; integrates with readiness checks and load balancing; accounts for startup and warm-up time; and uses separate scale-out and scale-in behaviour to prevent oscillation. Scale out quickly enough to protect users, but scale in cautiously and drain work before termination. Autoscaling cannot correct inefficient queries, hot partitions, memory leaks, or exhausted downstream services. Budget database connections and external concurrency at maximum scale, maintain failure headroom, apply rate limiting while additional capacity starts, and alert when scaling reaches its configured limit. Finally, validate the complete control loop under realistic traffic, failures, deployments, and dependency constraints.