health checks
Health Checks
Learn how liveness, readiness, startup, dependency, and external health checks help load balancers and orchestrators decide whether an application instance should receive traffic, remain running, restart, drain connections, or trigger an operational alert.
Introduction
A process can be running without being capable of serving a valid request. For example, an application can remain visible in the operating-system process list while its worker threads are blocked, initialization is incomplete, required configuration is missing, or database connections are exhausted.
Load balancers and orchestration platforms need a reliable way to determine whether an application instance should receive traffic or be restarted. Health checks provide this information through automated probes.
A health check sends a controlled request or operation to an application instance and interprets the result according to a defined contract.
Core idea: Different health checks control different actions. Liveness determines whether a process should restart. Readiness determines whether an instance should receive traffic. Startup determines whether initialization has completed.
Application Instance
|
+-- Startup check
| Has initialization completed?
|
+-- Liveness check
| Is the process capable of making progress?
|
+-- Readiness check
Can this instance safely receive traffic?
Treating all these questions as one generic health check can create self-inflicted outages. A temporary database problem should not necessarily restart every application process, and an application that is still starting should not receive production traffic.
Prerequisites
| # | Prerequisite | Why It Is Needed |
|---|---|---|
| 1 | Horizontal scaling | Health checks determine which instances belong to the active backend pool. |
| 2 | Load balancing | A load balancer uses health information to avoid unavailable backends. |
| 3 | Reverse proxies and API gateways | These components commonly route traffic only to ready backends. |
| 4 | HTTP response codes | HTTP health endpoints communicate healthy and unavailable states. |
| 5 | Application dependencies | Databases, caches, queues, and external APIs can influence readiness. |
| 6 | Observability | Metrics and logs help distinguish a real failure from a faulty probe. |
What Is a Health Check?
A health check is an automated test used to evaluate the operational state of a service, process, server, container, or dependency.
Health-check system
|
v
Send probe to application instance
|
+-- Successful result:
| take configured healthy action
|
+-- Failed result:
take configured unhealthy action
The action depends on the health-check type. The system can:
- Add an instance to traffic rotation
- Remove an instance from traffic rotation
- Restart a process or container
- Delay traffic during application startup
- Trigger replacement of an unhealthy instance
- Raise an operational alert
- Initiate failover
Health Check vs Monitoring
| Area | Health Check | Monitoring |
|---|---|---|
| Primary question | Should this instance run or receive traffic? | How is the system behaving over time? |
| Typical result | Healthy, not ready, or unhealthy | Metrics, logs, traces, trends, and alerts |
| Typical consumer | Load balancer, orchestrator, or repair controller | Operators, dashboards, and alerting systems |
| Typical action | Route, remove, restart, or replace | Investigate, alert, scale, or optimize |
| Detail level | Focused operational verdict | Detailed diagnostic and performance information |
Health checks should not replace metrics and diagnostic monitoring. A healthy result does not prove that latency, error rate, throughput, or cost meets the required objective.
Main Health-check Types
| Check | Question | Typical Failure Action |
|---|---|---|
| Startup | Has the application completed initialization? | Continue waiting or restart after the configured startup boundary |
| Liveness | Is the process capable of making progress? | Restart the process or container |
| Readiness | Can the instance safely serve new traffic? | Remove the instance from active routing |
| Dependency | Is a required downstream dependency usable? | Degrade readiness, raise an alert, or use a fallback |
| External or synthetic | Can a user-visible journey succeed from outside? | Raise an alert or initiate failover according to policy |
Startup Check
A startup check determines whether an application has finished its initialization process.
Process starts
|
v
Load configuration
|
v
Initialize runtime
|
v
Load required assets
|
v
Create initial connections
|
v
Warm required components
|
v
Startup check succeeds
Until startup succeeds, the platform can delay liveness and readiness decisions that would otherwise restart or route traffic to the application prematurely.
Suitable Startup Conditions
- Required configuration loaded
- Application container initialized
- Required migrations or startup validation completed
- Runtime is capable of accepting probe requests
- Required local assets are available
Startup rule: Use startup checks for initialization that can take longer than ordinary health-check timing. Do not make liveness kill a correctly starting application repeatedly.
Liveness Check
A liveness check answers whether the process remains capable of making progress.
Liveness succeeds:
Process event loop progresses
Worker thread can execute
Internal watchdog updates
Liveness fails:
Process deadlocked
Event loop permanently blocked
Critical internal runtime failure
A liveness failure commonly results in a process or container restart.
Good Liveness Characteristics
- Fast
- Local to the process
- Independent of external services
- Stable during temporary downstream outages
- Focused on failures that a restart can correct
Poor Liveness Dependencies
- External payment provider
- Remote search service
- Database query
- Shared cache availability
- Message-broker connection
- Internet connectivity
If a shared database fails and every application's liveness check fails, the orchestrator can restart every application instance. The restarts do not repair the database and can add further load during recovery.
Liveness rule: Fail liveness only when restarting the process is a reasonable corrective action. Do not restart healthy processes merely because a downstream service is unavailable.
Readiness Check
A readiness check answers whether the instance can safely receive new production traffic.
Load Balancer
|
v
Readiness probe
|
+-- Ready:
| include instance
| in backend pool
|
+-- Not ready:
stop sending
new requests
Readiness can fail temporarily without requiring a restart.
Possible Readiness Conditions
- Startup completed
- Required configuration loaded
- The instance is not shutting down
- Required local resources are available
- Connection pools have usable capacity
- The instance is not critically overloaded
- Required dependencies are available according to policy
Liveness vs Readiness
| Area | Liveness | Readiness |
|---|---|---|
| Question | Should this process remain running? | Should this instance receive new traffic? |
| Failure action | Restart the process or container | Remove the instance from active routing |
| Dependency checks | Generally avoid external dependencies | Can include carefully selected required dependencies |
| Temporary overload | Should not normally trigger restart | Can temporarily remove the instance from rotation |
| Deployment draining | Process can remain alive | Marked not ready before termination |
Startup vs Readiness
| Area | Startup Check | Readiness Check |
|---|---|---|
| Purpose | Protect initialization | Control production traffic membership |
| Lifecycle | Primarily during application startup | Continues while the application runs |
| Failure meaning | Initialization is not complete | The instance cannot currently serve new traffic |
| Recovery | Continue starting or restart after an allowed boundary | Become ready again when the temporary condition clears |
L4 Health Checks
A Layer 4 health check can verify whether a TCP connection can be established to a configured backend port.
Health checker
|
v
Open TCP connection
to backend port
|
+-- Connection accepted:
| transport endpoint available
|
+-- Connection refused or timed out:
transport endpoint unavailable
A successful Layer 4 check proves that a process accepted the connection. It does not prove that the application can complete a useful request.
L7 Health Checks
A Layer 7 health check sends an application-level request and evaluates the response.
GET /health/ready HTTP/1.1
Host: api.example.com
User-Agent: platform-health-check
HTTP/1.1 200 OK
Content-Type: application/json
Cache-Control: no-store
{
"status": "ready"
}
An HTTP-based probe can detect an application that accepts transport connections but cannot process HTTP requests correctly.
L4 vs L7 Health Checks
| Area | L4 Probe | L7 Probe |
|---|---|---|
| Typical operation | Open a TCP connection | Send an application request |
| Application awareness | Limited | Application-aware |
| Cost | Generally simpler | Requires protocol handling |
| Detects open port | Yes | Yes, when the application request establishes a connection |
| Detects invalid application response | No | Can detect it |
Recommended Endpoint Separation
/health/startup
Question:
Has application initialization completed?
/health/live
Question:
Is this process capable of making progress?
/health/ready
Question:
Can this instance safely receive traffic?
/health/details
Purpose:
Protected diagnostic information for operators.
Endpoint names are design choices. The important requirement is that each endpoint has one clear purpose and a predictable operational consequence.
Public vs Protected Health Endpoints
A basic health endpoint can be reachable by infrastructure without exposing sensitive implementation details.
{
"status": "ready"
}
{
"databaseHost": "internal-db-name",
"databaseUser": "application-user",
"cacheHost": "internal-cache-name",
"connectionString": "sensitive-value",
"exception": "complete internal stack trace"
}
Detailed health diagnostics should be protected by appropriate authentication, authorization, network restrictions, and logging controls.
Dependency Checks
Applications often depend on databases, caches, queues, object storage, and other services.
Application
|
+-- Database
+-- Cache
+-- Message Broker
+-- Object Storage
+-- External Service
Not every dependency should be included in every health check.
Questions to Ask
- Is the dependency required for every request?
- Can the application serve a degraded response without it?
- Will failing readiness reduce or worsen the incident?
- Can all instances fail the same dependency check together?
- Can the check create meaningful load on the dependency?
- Is cached dependency status sufficient?
- Can a circuit breaker provide a safer signal?
Dependency rule: A deep dependency check can detect real failures, but it can also remove every application instance when one shared dependency fails. Evaluate the routing consequence before adding a dependency to readiness.
Required vs Optional Dependencies
| Dependency | Possible Classification | Possible Behaviour |
|---|---|---|
| Primary transactional database | Required for transactional APIs | Selected write routes may become not ready |
| Recommendation service | Optional for core course display | Serve the course without recommendations |
| Cache | Optional when the source database is available | Bypass cache with controlled protection |
| Email provider | Optional for the immediate request | Queue the notification for later processing |
| Object storage | Required only for media operations | Media routes can fail while unrelated routes remain available |
Classification depends on the specific service and operation. One global readiness verdict can be too coarse for an application with independent capabilities.
Shallow vs Deep Checks
Shallow Check
Verify:
- Process responds
- Internal runtime is progressing
- Required local configuration is loaded
- Instance is not draining
Deep Check
Verify:
- Database query
- Cache operation
- Queue operation
- Object-storage operation
- External service request
Shallow checks are generally better for frequent routing decisions. Deep checks can be useful for protected diagnostics, synthetic monitoring, and carefully designed readiness conditions.
Health-check Amplification
Health checks themselves create traffic.
If \(I\) instances are checked every \(T\) seconds, the approximate check rate is:
\[ HealthCheckRate = \frac{I}{T} \]
If several load-balancer nodes, orchestrators, and monitoring systems probe independently, actual traffic can be higher.
100 application instances
Each readiness check executes:
- 1 database query
- 1 cache command
- 1 external API request
Result:
Health-check traffic reaches and
depends on every shared dependency.
Avoid turning a health endpoint into a recurring load-test request.
Probe Timing
Health-check timing controls how quickly a failure is detected and how easily temporary problems create false removals or restart loops.
| Setting | Purpose |
|---|---|
| Initial delay | Wait before beginning checks after process startup |
| Probe interval | Time between checks |
| Probe timeout | Maximum time allowed for one check |
| Failure threshold | Number of failures required before taking the unhealthy action |
| Success threshold | Number of successes required before restoring healthy state |
| Grace period | Time permitted for shutdown or recovery behaviour |
Exact values should come from observed startup times, failure modes, transient latency, routing requirements, and recovery objectives.
Detection-time Intuition
A simplified failure-detection estimate is:
\[ DetectionTime \approx ProbeInterval \times FailureThreshold \]
Timeout duration, scheduling delay, network latency, and load-balancer implementation can increase the actual time.
More frequent checks detect failures faster but create more traffic and can react more aggressively to brief interruptions.
Preventing Health Flapping
Flapping occurs when an instance repeatedly moves between ready and not-ready states.
Probe succeeds
-> add instance
Probe fails briefly
-> remove instance
Next probe succeeds
-> add instance
Load increases
-> probe fails again
Possible controls include:
- Consecutive failure thresholds
- Consecutive success thresholds
- Cooldown periods
- Readiness hysteresis
- Cached dependency state
- Capacity-aware load balancing
- Removing the root cause of overload
Cascading Readiness Failure
Readiness failure can reduce capacity and increase load on the remaining instances.
One instance becomes overloaded
|
v
Readiness fails
|
v
Traffic moves to remaining instances
|
v
Remaining instances become overloaded
|
v
Their readiness checks fail
|
v
Service loses all healthy backends
Readiness should not be used as the only overload-control mechanism. Capacity planning, autoscaling, rate limiting, admission control, queueing, and load shedding can be safer controls.
Health Checks during Deployment
Start new instance
|
v
Startup check runs
|
v
Readiness succeeds
|
v
Instance enters traffic pool
|
v
Old instance marked not ready
|
v
New requests stop
|
v
In-flight requests drain
|
v
Old instance terminates
This sequence supports rolling deployment without sending traffic to an uninitialized instance or terminating an instance still processing work.
Graceful Shutdown and Readiness
When an instance begins shutdown, it should report not ready before the process exits.
Termination requested
|
v
Set readiness to false
|
v
Load balancer stops new routing
|
v
Drain in-flight work
|
v
Close background consumers
|
v
Close connections
|
v
Exit process
The time needed to update backend membership must be included in the shutdown grace period.
Automatic Repair
Infrastructure platforms can use per-instance health signals to replace instances that remain unhealthy.
Instance health repeatedly fails
|
v
Instance removed from traffic
|
v
Repair controller evaluates state
|
v
Unhealthy instance replaced
|
v
Replacement initializes
|
v
Readiness succeeds
|
v
Replacement joins backend pool
Automatic repair is useful when instances are disposable and initialization is reliable. It can create repeated replacement loops when the fault is in a shared configuration, image, dependency, or network.
Kubernetes Probe Example
apiVersion: apps/v1
kind: Deployment
metadata:
name: course-api
spec:
replicas: 3
template:
spec:
containers:
- name: course-api
image: approved-course-api-image
ports:
- containerPort: 8080
startupProbe:
httpGet:
path: /health/startup
port: 8080
livenessProbe:
httpGet:
path: /health/live
port: 8080
readinessProbe:
httpGet:
path: /health/ready
port: 8080
This example shows separation of probe purposes. Production timing, thresholds, security, termination, and resource settings must be selected from tested application behaviour.
PHP Health Controller
<?php
declare(strict_types=1);
final class HealthController
{
public function __construct(
private ApplicationState $applicationState,
private ShutdownState $shutdownState
) {
}
public function live(): void
{
$this->writeResponse(
200,
[
'status' => 'alive'
]
);
}
public function ready(): void
{
if ($this->shutdownState->isDraining()) {
$this->writeResponse(
503,
[
'status' => 'not-ready'
]
);
return;
}
if (!$this->applicationState->isInitialized()) {
$this->writeResponse(
503,
[
'status' => 'not-ready'
]
);
return;
}
$this->writeResponse(
200,
[
'status' => 'ready'
]
);
}
public function startup(): void
{
if (!$this->applicationState->isInitialized()) {
$this->writeResponse(
503,
[
'status' => 'starting'
]
);
return;
}
$this->writeResponse(
200,
[
'status' => 'started'
]
);
}
private function writeResponse(
int $statusCode,
array $payload
): void {
http_response_code(
$statusCode
);
header(
'Content-Type: application/json'
);
header(
'Cache-Control: no-store'
);
echo json_encode(
$payload,
JSON_THROW_ON_ERROR
);
}
}
This example keeps the public health response minimal. Production code should use your application's existing routing, response, error-handling, and dependency abstractions.
Protected Diagnostic Response
{
"status": "degraded",
"checks": {
"runtime": {
"status": "healthy"
},
"database": {
"status": "healthy"
},
"cache": {
"status": "unavailable",
"required": false
}
},
"serviceVersion": "approved-service-version"
}
This type of detailed response should not expose credentials, connection strings, internal addresses, stack traces, or confidential configuration.
Health-status Model
| Status | Meaning | Possible Routing Behaviour |
|---|---|---|
| Starting | Initialization is still in progress | Do not route production traffic |
| Healthy and ready | The instance can serve its intended traffic | Include in the backend pool |
| Degraded | Some optional capability is unavailable | Route only if the supported operation remains safe |
| Not ready | The process is alive but cannot serve new traffic safely | Remove from active routing |
| Unhealthy | The process cannot make progress or recover without intervention | Restart, replace, or alert according to policy |
| Draining | The instance is completing existing work before shutdown | Do not send new traffic |
Synthetic Health Checks
A synthetic check exercises a controlled user-visible journey from outside the application process.
Synthetic monitor
|
v
Resolve DNS
|
v
Establish TLS
|
v
Call API gateway or web endpoint
|
v
Authenticate using a controlled test identity
|
v
Execute a safe representative operation
|
v
Validate response
Synthetic checks can detect problems that an internal process check misses, including DNS, certificates, routing, gateway policy, and selected backend failures.
Synthetic operations must be safe, isolated, identifiable, and cleaned up where they create data.
Security Considerations
Protect health-check systems against:
- Exposure of internal topology
- Exposure of software versions
- Exposure of credentials or configuration
- Unauthenticated access to detailed diagnostics
- Expensive public probe abuse
- Cache poisoning
- Forged health-check headers
- Direct backend bypass
Load balancers and orchestrators should reach health endpoints through approved network paths. If a special health-check header is used, do not treat the presence of that header alone as proof of a trusted caller.
Do Not Cache Health Responses
Health endpoints should return the current instance state rather than an old cached response.
HTTP/1.1 200 OK
Content-Type: application/json
Cache-Control: no-store
A stale healthy result can keep an unavailable backend in traffic rotation. A stale unhealthy result can keep a recovered backend unavailable.
Observability
Useful health-check metrics include:
- Startup duration
- Liveness probe success and failure count
- Readiness probe success and failure count
- Probe latency
- Healthy backend count
- Backend removal count
- Backend restoration count
- Process restart count
- Automatic repair count
- Health-state transition count
- Dependency-check latency
- Time spent not ready
- Connection-draining duration
- Synthetic-check success rate
Alert Conditions
Alert when:
- Healthy backend count falls below the required minimum
- All instances become not ready
- An instance repeatedly restarts
- Startup repeatedly exceeds its expected operational boundary
- Readiness rapidly changes between states
- Probe latency increases unexpectedly
- A shared dependency causes widespread readiness failure
- Automatic instance replacement repeats without recovery
- Synthetic checks fail while internal probes remain healthy
- An instance remains draining beyond the allowed period
Troubleshooting Workflow
- Identify which probe is failing: startup, liveness, or readiness.
- Identify what action the failed probe triggers.
- Call the health endpoint through the expected network path.
- Check response status, body, and latency.
- Check application logs using the probe request time.
- Check whether the application completed startup.
- Check whether the instance is intentionally draining.
- Check local CPU, memory, threads, file descriptors, and connections.
- Check required dependency state.
- Check probe interval, timeout, and thresholds.
- Check load-balancer or orchestrator backend membership.
- Check network rules between the checker and the application.
- Check recent application or infrastructure deployments.
- Determine whether restart, traffic removal, degradation, or alerting is the correct action.
Common Health-check Mistakes
Using One Endpoint for Every Decision
Startup, restart, routing, diagnostics, and external availability require different questions and actions.
Checking External Dependencies in Liveness
A shared dependency outage can restart every healthy application process without correcting the actual failure.
Sending Traffic before Readiness
Requests reach an instance before initialization, configuration, or warm-up is complete.
Using an Expensive Database Query
Frequent health probes add load and can contribute to the dependency failure they are intended to detect.
Failing Every Instance on One Shared Dependency
The load balancer can lose all eligible backends even though a degraded application response remains possible.
Using Very Aggressive Probe Timing
Short delays and low thresholds can produce false failures and restart loops during normal latency variation.
Using Very Slow Detection
Unavailable instances remain in the backend pool and continue receiving traffic.
Returning HTTP 200 for Every State
Routing infrastructure cannot distinguish ready from unavailable instances when every response appears successful.
Caching Health Responses
Infrastructure can receive a stale healthy or stale unhealthy result.
Exposing Detailed Diagnostics Publicly
Internal topology, software details, exceptions, and dependency information can be disclosed.
Removing an Instance before Draining
Active requests can be interrupted if readiness and connection draining are not coordinated.
Assuming Healthy Means Fast
A health probe can succeed while user requests remain slow, error-prone, or unable to meet the service objective.
Recommended Test Cases
| Test | Expected Evidence |
|---|---|
| Normal startup | No production traffic is routed before initialization completes |
| Slow startup | The startup probe prevents premature liveness restarts |
| Internal process deadlock | Liveness eventually fails and the process is restarted |
| Temporary not-ready state | The instance leaves routing without being restarted unnecessarily |
| Readiness recovery | The instance rejoins routing after the required success condition |
| Database outage | Liveness remains independent and readiness follows the documented dependency policy |
| Optional cache outage | The service follows its documented degraded behaviour |
| Probe timeout | The checker applies the configured failure threshold |
| Transient probe failure | One brief failure does not create unnecessary flapping |
| Deployment startup | A new instance becomes ready before an old instance drains |
| Graceful shutdown | No new requests arrive after readiness becomes false |
| Automatic repair | A persistently unhealthy disposable instance is replaced |
| Detailed endpoint security | Unauthorized users cannot access internal diagnostic information |
| Health-response caching | Infrastructure always receives the current health result |
| Synthetic journey | DNS, TLS, routing, gateway, and backend behaviour are tested externally |
Health-check Best Practices
Recommended Practices
- Create separate startup, liveness, and readiness checks.
- Define the action controlled by every probe.
- Keep liveness local to the process.
- Fail liveness only when restarting can reasonably help.
- Use readiness to control traffic membership.
- Mark an instance not ready before graceful shutdown.
- Use startup checks for applications with long initialization.
- Keep frequently executed checks fast and bounded.
- Avoid expensive queries in health endpoints.
- Classify dependencies as required or optional.
- Consider the cluster-wide result of a shared dependency failure.
- Use consecutive failure and success thresholds.
- Tune probe timeouts using representative measurements.
- Prevent cached health responses.
- Return minimal public health information.
- Protect detailed diagnostic endpoints.
- Monitor health-state changes and restart loops.
- Use synthetic checks for external user-visible journeys.
- Test health behaviour during deployments and dependency failures.
- Do not equate a successful health check with complete service quality.
Practice Exercise
Design health checks for your online learning platform's course API.
Requirements
- Create separate startup, liveness, and readiness endpoints.
- Define the exact question answered by each endpoint.
- Keep liveness independent of the database.
- Prevent traffic before configuration loading is complete.
- Define database behaviour for course-read operations.
- Define degraded behaviour when the recommendation service is unavailable.
- Define readiness during connection-pool exhaustion.
- Prevent public access to detailed diagnostics.
- Add no-cache response headers.
- Define probe interval, timeout, and thresholds from test results.
- Mark the instance not ready during shutdown.
- Drain existing requests before termination.
- Test a slow startup.
- Test a database outage.
- Test an internal application deadlock.
- Add an external synthetic course-read test.
- Monitor startup duration, readiness changes, and restarts.
Health-check Design Template
| Check | Question | Included Conditions | Failure Action |
|---|---|---|---|
| Startup | Has initialization completed? | Configuration and required local initialization | Continue waiting or restart after the approved boundary |
| Liveness | Can the process make progress? | Internal runtime or watchdog state | Restart the process |
| Readiness | Can the instance serve new API requests? | Initialization, draining state, and essential serving capacity | Remove from active routing |
| Detailed diagnostics | Which component is degraded? | Protected dependency and runtime information | Support investigation and alerts |
| Synthetic check | Can a controlled user journey succeed externally? | DNS, TLS, gateway, authentication, routing, and API response | Raise an alert or perform approved failover |
Frequently Asked Questions
What is a health check?
A health check is an automated probe used to determine whether a process, instance, service, or dependency is in a state suitable for a defined operational action.
What is a liveness check?
A liveness check determines whether the application process can continue running and making progress.
What happens when liveness fails?
The hosting or orchestration platform commonly restarts the failed process or container according to its configured policy.
What is a readiness check?
A readiness check determines whether an instance can safely receive new production traffic.
What happens when readiness fails?
The instance is normally removed from active routing without necessarily restarting the process.
What is a startup check?
A startup check determines whether application initialization has completed before ordinary liveness and readiness decisions begin.
Should liveness check the database?
Generally no. A database outage does not normally mean that restarting the application process will correct the problem.
Should readiness check the database?
It depends on whether the database is required for the traffic handled by that instance and whether removing all instances would improve or worsen the failure.
What is the difference between an L4 and L7 health check?
An L4 health check normally verifies transport connectivity. An L7 health check sends an application-level request and evaluates the response.
Should health endpoints return detailed diagnostics?
Infrastructure-facing endpoints should normally return a minimal verdict. Detailed diagnostics should be protected from unauthorized access.
Can a health check cause an outage?
Yes. An overly aggressive or dependency-heavy check can restart healthy instances, remove every backend, amplify dependency traffic, or create repeated health flapping.
Does a healthy response prove that the service is performing well?
No. Health indicates suitability for a specific operational decision. Latency, throughput, errors, resource utilization, and user journeys still require monitoring.
Key Takeaway
Health checks are operational control signals, not general diagnostic pages. A startup check protects initialization, a liveness check determines whether a process should restart, and a readiness check determines whether an instance should receive new traffic. Keep liveness local and fail it only when a restart can help. Design readiness around the service's actual ability to serve traffic, while considering whether shared dependency failures could remove every backend. Keep frequent checks fast, bounded, and uncached, use thresholds to prevent flapping, and protect detailed diagnostics. Coordinate readiness with deployments and graceful shutdown so new instances receive traffic only after initialization and old instances drain before termination. Finally, combine internal probes with metrics, tracing, alerts, and external synthetic checks to verify the complete user-visible service.