Table of Contents

    health checks

    LOAD BALANCING, PROXIES & ELASTIC SCALING

    Health Checks

    Learn how liveness, readiness, startup, dependency, and external health checks help load balancers and orchestrators decide whether an application instance should receive traffic, remain running, restart, drain connections, or trigger an operational alert.

    Introduction

    A process can be running without being capable of serving a valid request. For example, an application can remain visible in the operating-system process list while its worker threads are blocked, initialization is incomplete, required configuration is missing, or database connections are exhausted.

    Load balancers and orchestration platforms need a reliable way to determine whether an application instance should receive traffic or be restarted. Health checks provide this information through automated probes.

    A health check sends a controlled request or operation to an application instance and interprets the result according to a defined contract.

    Core idea: Different health checks control different actions. Liveness determines whether a process should restart. Readiness determines whether an instance should receive traffic. Startup determines whether initialization has completed.

    Application Instance
            |
            +-- Startup check
            |      Has initialization completed?
            |
            +-- Liveness check
            |      Is the process capable of making progress?
            |
            +-- Readiness check
                   Can this instance safely receive traffic?

    Treating all these questions as one generic health check can create self-inflicted outages. A temporary database problem should not necessarily restart every application process, and an application that is still starting should not receive production traffic.

    Prerequisites

    # Prerequisite Why It Is Needed
    1 Horizontal scaling Health checks determine which instances belong to the active backend pool.
    2 Load balancing A load balancer uses health information to avoid unavailable backends.
    3 Reverse proxies and API gateways These components commonly route traffic only to ready backends.
    4 HTTP response codes HTTP health endpoints communicate healthy and unavailable states.
    5 Application dependencies Databases, caches, queues, and external APIs can influence readiness.
    6 Observability Metrics and logs help distinguish a real failure from a faulty probe.

    What Is a Health Check?

    A health check is an automated test used to evaluate the operational state of a service, process, server, container, or dependency.

    Health-check system
            |
            v
    Send probe to application instance
            |
            +-- Successful result:
            |      take configured healthy action
            |
            +-- Failed result:
                   take configured unhealthy action

    The action depends on the health-check type. The system can:

    • Add an instance to traffic rotation
    • Remove an instance from traffic rotation
    • Restart a process or container
    • Delay traffic during application startup
    • Trigger replacement of an unhealthy instance
    • Raise an operational alert
    • Initiate failover
    Health-check Design Flow
    define the question → define the probe → define success → define failure threshold → define the resulting action

    Health Check vs Monitoring

    Area Health Check Monitoring
    Primary question Should this instance run or receive traffic? How is the system behaving over time?
    Typical result Healthy, not ready, or unhealthy Metrics, logs, traces, trends, and alerts
    Typical consumer Load balancer, orchestrator, or repair controller Operators, dashboards, and alerting systems
    Typical action Route, remove, restart, or replace Investigate, alert, scale, or optimize
    Detail level Focused operational verdict Detailed diagnostic and performance information

    Health checks should not replace metrics and diagnostic monitoring. A healthy result does not prove that latency, error rate, throughput, or cost meets the required objective.

    Main Health-check Types

    Check Question Typical Failure Action
    Startup Has the application completed initialization? Continue waiting or restart after the configured startup boundary
    Liveness Is the process capable of making progress? Restart the process or container
    Readiness Can the instance safely serve new traffic? Remove the instance from active routing
    Dependency Is a required downstream dependency usable? Degrade readiness, raise an alert, or use a fallback
    External or synthetic Can a user-visible journey succeed from outside? Raise an alert or initiate failover according to policy

    Startup Check

    A startup check determines whether an application has finished its initialization process.

    Process starts
          |
          v
    Load configuration
          |
          v
    Initialize runtime
          |
          v
    Load required assets
          |
          v
    Create initial connections
          |
          v
    Warm required components
          |
          v
    Startup check succeeds

    Until startup succeeds, the platform can delay liveness and readiness decisions that would otherwise restart or route traffic to the application prematurely.

    Suitable Startup Conditions

    • Required configuration loaded
    • Application container initialized
    • Required migrations or startup validation completed
    • Runtime is capable of accepting probe requests
    • Required local assets are available

    Startup rule: Use startup checks for initialization that can take longer than ordinary health-check timing. Do not make liveness kill a correctly starting application repeatedly.

    Liveness Check

    A liveness check answers whether the process remains capable of making progress.

    Liveness succeeds:
    
    Process event loop progresses
    Worker thread can execute
    Internal watchdog updates
    
    
    Liveness fails:
    
    Process deadlocked
    Event loop permanently blocked
    Critical internal runtime failure

    A liveness failure commonly results in a process or container restart.

    Good Liveness Characteristics

    • Fast
    • Local to the process
    • Independent of external services
    • Stable during temporary downstream outages
    • Focused on failures that a restart can correct

    Poor Liveness Dependencies

    • External payment provider
    • Remote search service
    • Database query
    • Shared cache availability
    • Message-broker connection
    • Internet connectivity

    If a shared database fails and every application's liveness check fails, the orchestrator can restart every application instance. The restarts do not repair the database and can add further load during recovery.

    Liveness rule: Fail liveness only when restarting the process is a reasonable corrective action. Do not restart healthy processes merely because a downstream service is unavailable.

    Readiness Check

    A readiness check answers whether the instance can safely receive new production traffic.

    Load Balancer
          |
          v
    Readiness probe
          |
          +-- Ready:
          |      include instance
          |      in backend pool
          |
          +-- Not ready:
                 stop sending
                 new requests

    Readiness can fail temporarily without requiring a restart.

    Possible Readiness Conditions

    • Startup completed
    • Required configuration loaded
    • The instance is not shutting down
    • Required local resources are available
    • Connection pools have usable capacity
    • The instance is not critically overloaded
    • Required dependencies are available according to policy

    Liveness vs Readiness

    Area Liveness Readiness
    Question Should this process remain running? Should this instance receive new traffic?
    Failure action Restart the process or container Remove the instance from active routing
    Dependency checks Generally avoid external dependencies Can include carefully selected required dependencies
    Temporary overload Should not normally trigger restart Can temporarily remove the instance from rotation
    Deployment draining Process can remain alive Marked not ready before termination

    Startup vs Readiness

    Area Startup Check Readiness Check
    Purpose Protect initialization Control production traffic membership
    Lifecycle Primarily during application startup Continues while the application runs
    Failure meaning Initialization is not complete The instance cannot currently serve new traffic
    Recovery Continue starting or restart after an allowed boundary Become ready again when the temporary condition clears

    L4 Health Checks

    A Layer 4 health check can verify whether a TCP connection can be established to a configured backend port.

    Health checker
          |
          v
    Open TCP connection
    to backend port
          |
          +-- Connection accepted:
          |      transport endpoint available
          |
          +-- Connection refused or timed out:
                 transport endpoint unavailable

    A successful Layer 4 check proves that a process accepted the connection. It does not prove that the application can complete a useful request.

    L7 Health Checks

    A Layer 7 health check sends an application-level request and evaluates the response.

    GET /health/ready HTTP/1.1
    Host: api.example.com
    User-Agent: platform-health-check
    HTTP/1.1 200 OK
    Content-Type: application/json
    Cache-Control: no-store
    {
      "status": "ready"
    }

    An HTTP-based probe can detect an application that accepts transport connections but cannot process HTTP requests correctly.

    L4 vs L7 Health Checks

    Area L4 Probe L7 Probe
    Typical operation Open a TCP connection Send an application request
    Application awareness Limited Application-aware
    Cost Generally simpler Requires protocol handling
    Detects open port Yes Yes, when the application request establishes a connection
    Detects invalid application response No Can detect it

    Recommended Endpoint Separation

    /health/startup
    
    Question:
    Has application initialization completed?
    
    
    /health/live
    
    Question:
    Is this process capable of making progress?
    
    
    /health/ready
    
    Question:
    Can this instance safely receive traffic?
    
    
    /health/details
    
    Purpose:
    Protected diagnostic information for operators.

    Endpoint names are design choices. The important requirement is that each endpoint has one clear purpose and a predictable operational consequence.

    Public vs Protected Health Endpoints

    A basic health endpoint can be reachable by infrastructure without exposing sensitive implementation details.

    Minimal infrastructure response
    {
      "status": "ready"
    }
    Excessive public detail
    {
      "databaseHost": "internal-db-name",
      "databaseUser": "application-user",
      "cacheHost": "internal-cache-name",
      "connectionString": "sensitive-value",
      "exception": "complete internal stack trace"
    }

    Detailed health diagnostics should be protected by appropriate authentication, authorization, network restrictions, and logging controls.

    Dependency Checks

    Applications often depend on databases, caches, queues, object storage, and other services.

    Application
        |
        +-- Database
        +-- Cache
        +-- Message Broker
        +-- Object Storage
        +-- External Service

    Not every dependency should be included in every health check.

    Questions to Ask

    • Is the dependency required for every request?
    • Can the application serve a degraded response without it?
    • Will failing readiness reduce or worsen the incident?
    • Can all instances fail the same dependency check together?
    • Can the check create meaningful load on the dependency?
    • Is cached dependency status sufficient?
    • Can a circuit breaker provide a safer signal?

    Dependency rule: A deep dependency check can detect real failures, but it can also remove every application instance when one shared dependency fails. Evaluate the routing consequence before adding a dependency to readiness.

    Required vs Optional Dependencies

    Dependency Possible Classification Possible Behaviour
    Primary transactional database Required for transactional APIs Selected write routes may become not ready
    Recommendation service Optional for core course display Serve the course without recommendations
    Cache Optional when the source database is available Bypass cache with controlled protection
    Email provider Optional for the immediate request Queue the notification for later processing
    Object storage Required only for media operations Media routes can fail while unrelated routes remain available

    Classification depends on the specific service and operation. One global readiness verdict can be too coarse for an application with independent capabilities.

    Shallow vs Deep Checks

    Shallow Check

    Verify:
    
    - Process responds
    - Internal runtime is progressing
    - Required local configuration is loaded
    - Instance is not draining

    Deep Check

    Verify:
    
    - Database query
    - Cache operation
    - Queue operation
    - Object-storage operation
    - External service request

    Shallow checks are generally better for frequent routing decisions. Deep checks can be useful for protected diagnostics, synthetic monitoring, and carefully designed readiness conditions.

    Health-check Amplification

    Health checks themselves create traffic.

    If \(I\) instances are checked every \(T\) seconds, the approximate check rate is:

    \[ HealthCheckRate = \frac{I}{T} \]

    If several load-balancer nodes, orchestrators, and monitoring systems probe independently, actual traffic can be higher.

    100 application instances
    
    Each readiness check executes:
    
    - 1 database query
    - 1 cache command
    - 1 external API request
    
    
    Result:
    
    Health-check traffic reaches and
    depends on every shared dependency.

    Avoid turning a health endpoint into a recurring load-test request.

    Probe Timing

    Health-check timing controls how quickly a failure is detected and how easily temporary problems create false removals or restart loops.

    Setting Purpose
    Initial delay Wait before beginning checks after process startup
    Probe interval Time between checks
    Probe timeout Maximum time allowed for one check
    Failure threshold Number of failures required before taking the unhealthy action
    Success threshold Number of successes required before restoring healthy state
    Grace period Time permitted for shutdown or recovery behaviour

    Exact values should come from observed startup times, failure modes, transient latency, routing requirements, and recovery objectives.

    Detection-time Intuition

    A simplified failure-detection estimate is:

    \[ DetectionTime \approx ProbeInterval \times FailureThreshold \]

    Timeout duration, scheduling delay, network latency, and load-balancer implementation can increase the actual time.

    More frequent checks detect failures faster but create more traffic and can react more aggressively to brief interruptions.

    Preventing Health Flapping

    Flapping occurs when an instance repeatedly moves between ready and not-ready states.

    Probe succeeds
        -> add instance
    
    Probe fails briefly
        -> remove instance
    
    Next probe succeeds
        -> add instance
    
    Load increases
        -> probe fails again

    Possible controls include:

    • Consecutive failure thresholds
    • Consecutive success thresholds
    • Cooldown periods
    • Readiness hysteresis
    • Cached dependency state
    • Capacity-aware load balancing
    • Removing the root cause of overload

    Cascading Readiness Failure

    Readiness failure can reduce capacity and increase load on the remaining instances.

    One instance becomes overloaded
          |
          v
    Readiness fails
          |
          v
    Traffic moves to remaining instances
          |
          v
    Remaining instances become overloaded
          |
          v
    Their readiness checks fail
          |
          v
    Service loses all healthy backends

    Readiness should not be used as the only overload-control mechanism. Capacity planning, autoscaling, rate limiting, admission control, queueing, and load shedding can be safer controls.

    Health Checks during Deployment

    Start new instance
          |
          v
    Startup check runs
          |
          v
    Readiness succeeds
          |
          v
    Instance enters traffic pool
          |
          v
    Old instance marked not ready
          |
          v
    New requests stop
          |
          v
    In-flight requests drain
          |
          v
    Old instance terminates

    This sequence supports rolling deployment without sending traffic to an uninitialized instance or terminating an instance still processing work.

    Graceful Shutdown and Readiness

    When an instance begins shutdown, it should report not ready before the process exits.

    Termination requested
          |
          v
    Set readiness to false
          |
          v
    Load balancer stops new routing
          |
          v
    Drain in-flight work
          |
          v
    Close background consumers
          |
          v
    Close connections
          |
          v
    Exit process

    The time needed to update backend membership must be included in the shutdown grace period.

    Automatic Repair

    Infrastructure platforms can use per-instance health signals to replace instances that remain unhealthy.

    Instance health repeatedly fails
          |
          v
    Instance removed from traffic
          |
          v
    Repair controller evaluates state
          |
          v
    Unhealthy instance replaced
          |
          v
    Replacement initializes
          |
          v
    Readiness succeeds
          |
          v
    Replacement joins backend pool

    Automatic repair is useful when instances are disposable and initialization is reliable. It can create repeated replacement loops when the fault is in a shared configuration, image, dependency, or network.

    Kubernetes Probe Example

    apiVersion: apps/v1
    kind: Deployment
    metadata:
      name: course-api
    spec:
      replicas: 3
    
      template:
        spec:
          containers:
            - name: course-api
              image: approved-course-api-image
    
              ports:
                - containerPort: 8080
    
              startupProbe:
                httpGet:
                  path: /health/startup
                  port: 8080
    
              livenessProbe:
                httpGet:
                  path: /health/live
                  port: 8080
    
              readinessProbe:
                httpGet:
                  path: /health/ready
                  port: 8080

    This example shows separation of probe purposes. Production timing, thresholds, security, termination, and resource settings must be selected from tested application behaviour.

    PHP Health Controller

    <?php
    
    declare(strict_types=1);
    
    final class HealthController
    {
        public function __construct(
            private ApplicationState $applicationState,
            private ShutdownState $shutdownState
        ) {
        }
    
        public function live(): void
        {
            $this->writeResponse(
                200,
                [
                    'status' => 'alive'
                ]
            );
        }
    
        public function ready(): void
        {
            if ($this->shutdownState->isDraining()) {
                $this->writeResponse(
                    503,
                    [
                        'status' => 'not-ready'
                    ]
                );
    
                return;
            }
    
            if (!$this->applicationState->isInitialized()) {
                $this->writeResponse(
                    503,
                    [
                        'status' => 'not-ready'
                    ]
                );
    
                return;
            }
    
            $this->writeResponse(
                200,
                [
                    'status' => 'ready'
                ]
            );
        }
    
        public function startup(): void
        {
            if (!$this->applicationState->isInitialized()) {
                $this->writeResponse(
                    503,
                    [
                        'status' => 'starting'
                    ]
                );
    
                return;
            }
    
            $this->writeResponse(
                200,
                [
                    'status' => 'started'
                ]
            );
        }
    
        private function writeResponse(
            int $statusCode,
            array $payload
        ): void {
            http_response_code(
                $statusCode
            );
    
            header(
                'Content-Type: application/json'
            );
    
            header(
                'Cache-Control: no-store'
            );
    
            echo json_encode(
                $payload,
                JSON_THROW_ON_ERROR
            );
        }
    }

    This example keeps the public health response minimal. Production code should use your application's existing routing, response, error-handling, and dependency abstractions.

    Protected Diagnostic Response

    {
      "status": "degraded",
      "checks": {
        "runtime": {
          "status": "healthy"
        },
        "database": {
          "status": "healthy"
        },
        "cache": {
          "status": "unavailable",
          "required": false
        }
      },
      "serviceVersion": "approved-service-version"
    }

    This type of detailed response should not expose credentials, connection strings, internal addresses, stack traces, or confidential configuration.

    Health-status Model

    Status Meaning Possible Routing Behaviour
    Starting Initialization is still in progress Do not route production traffic
    Healthy and ready The instance can serve its intended traffic Include in the backend pool
    Degraded Some optional capability is unavailable Route only if the supported operation remains safe
    Not ready The process is alive but cannot serve new traffic safely Remove from active routing
    Unhealthy The process cannot make progress or recover without intervention Restart, replace, or alert according to policy
    Draining The instance is completing existing work before shutdown Do not send new traffic

    Synthetic Health Checks

    A synthetic check exercises a controlled user-visible journey from outside the application process.

    Synthetic monitor
          |
          v
    Resolve DNS
          |
          v
    Establish TLS
          |
          v
    Call API gateway or web endpoint
          |
          v
    Authenticate using a controlled test identity
          |
          v
    Execute a safe representative operation
          |
          v
    Validate response

    Synthetic checks can detect problems that an internal process check misses, including DNS, certificates, routing, gateway policy, and selected backend failures.

    Synthetic operations must be safe, isolated, identifiable, and cleaned up where they create data.

    Security Considerations

    Protect health-check systems against:

    • Exposure of internal topology
    • Exposure of software versions
    • Exposure of credentials or configuration
    • Unauthenticated access to detailed diagnostics
    • Expensive public probe abuse
    • Cache poisoning
    • Forged health-check headers
    • Direct backend bypass

    Load balancers and orchestrators should reach health endpoints through approved network paths. If a special health-check header is used, do not treat the presence of that header alone as proof of a trusted caller.

    Do Not Cache Health Responses

    Health endpoints should return the current instance state rather than an old cached response.

    HTTP/1.1 200 OK
    Content-Type: application/json
    Cache-Control: no-store

    A stale healthy result can keep an unavailable backend in traffic rotation. A stale unhealthy result can keep a recovered backend unavailable.

    Observability

    Useful health-check metrics include:

    • Startup duration
    • Liveness probe success and failure count
    • Readiness probe success and failure count
    • Probe latency
    • Healthy backend count
    • Backend removal count
    • Backend restoration count
    • Process restart count
    • Automatic repair count
    • Health-state transition count
    • Dependency-check latency
    • Time spent not ready
    • Connection-draining duration
    • Synthetic-check success rate

    Alert Conditions

    Alert when:

    • Healthy backend count falls below the required minimum
    • All instances become not ready
    • An instance repeatedly restarts
    • Startup repeatedly exceeds its expected operational boundary
    • Readiness rapidly changes between states
    • Probe latency increases unexpectedly
    • A shared dependency causes widespread readiness failure
    • Automatic instance replacement repeats without recovery
    • Synthetic checks fail while internal probes remain healthy
    • An instance remains draining beyond the allowed period

    Troubleshooting Workflow

    1. Identify which probe is failing: startup, liveness, or readiness.
    2. Identify what action the failed probe triggers.
    3. Call the health endpoint through the expected network path.
    4. Check response status, body, and latency.
    5. Check application logs using the probe request time.
    6. Check whether the application completed startup.
    7. Check whether the instance is intentionally draining.
    8. Check local CPU, memory, threads, file descriptors, and connections.
    9. Check required dependency state.
    10. Check probe interval, timeout, and thresholds.
    11. Check load-balancer or orchestrator backend membership.
    12. Check network rules between the checker and the application.
    13. Check recent application or infrastructure deployments.
    14. Determine whether restart, traffic removal, degradation, or alerting is the correct action.

    Common Health-check Mistakes

    1

    Using One Endpoint for Every Decision

    Startup, restart, routing, diagnostics, and external availability require different questions and actions.

    2

    Checking External Dependencies in Liveness

    A shared dependency outage can restart every healthy application process without correcting the actual failure.

    3

    Sending Traffic before Readiness

    Requests reach an instance before initialization, configuration, or warm-up is complete.

    4

    Using an Expensive Database Query

    Frequent health probes add load and can contribute to the dependency failure they are intended to detect.

    5

    Failing Every Instance on One Shared Dependency

    The load balancer can lose all eligible backends even though a degraded application response remains possible.

    6

    Using Very Aggressive Probe Timing

    Short delays and low thresholds can produce false failures and restart loops during normal latency variation.

    7

    Using Very Slow Detection

    Unavailable instances remain in the backend pool and continue receiving traffic.

    8

    Returning HTTP 200 for Every State

    Routing infrastructure cannot distinguish ready from unavailable instances when every response appears successful.

    9

    Caching Health Responses

    Infrastructure can receive a stale healthy or stale unhealthy result.

    10

    Exposing Detailed Diagnostics Publicly

    Internal topology, software details, exceptions, and dependency information can be disclosed.

    11

    Removing an Instance before Draining

    Active requests can be interrupted if readiness and connection draining are not coordinated.

    12

    Assuming Healthy Means Fast

    A health probe can succeed while user requests remain slow, error-prone, or unable to meet the service objective.

    Recommended Test Cases

    Test Expected Evidence
    Normal startup No production traffic is routed before initialization completes
    Slow startup The startup probe prevents premature liveness restarts
    Internal process deadlock Liveness eventually fails and the process is restarted
    Temporary not-ready state The instance leaves routing without being restarted unnecessarily
    Readiness recovery The instance rejoins routing after the required success condition
    Database outage Liveness remains independent and readiness follows the documented dependency policy
    Optional cache outage The service follows its documented degraded behaviour
    Probe timeout The checker applies the configured failure threshold
    Transient probe failure One brief failure does not create unnecessary flapping
    Deployment startup A new instance becomes ready before an old instance drains
    Graceful shutdown No new requests arrive after readiness becomes false
    Automatic repair A persistently unhealthy disposable instance is replaced
    Detailed endpoint security Unauthorized users cannot access internal diagnostic information
    Health-response caching Infrastructure always receives the current health result
    Synthetic journey DNS, TLS, routing, gateway, and backend behaviour are tested externally

    Health-check Best Practices

    Recommended Practices

    • Create separate startup, liveness, and readiness checks.
    • Define the action controlled by every probe.
    • Keep liveness local to the process.
    • Fail liveness only when restarting can reasonably help.
    • Use readiness to control traffic membership.
    • Mark an instance not ready before graceful shutdown.
    • Use startup checks for applications with long initialization.
    • Keep frequently executed checks fast and bounded.
    • Avoid expensive queries in health endpoints.
    • Classify dependencies as required or optional.
    • Consider the cluster-wide result of a shared dependency failure.
    • Use consecutive failure and success thresholds.
    • Tune probe timeouts using representative measurements.
    • Prevent cached health responses.
    • Return minimal public health information.
    • Protect detailed diagnostic endpoints.
    • Monitor health-state changes and restart loops.
    • Use synthetic checks for external user-visible journeys.
    • Test health behaviour during deployments and dependency failures.
    • Do not equate a successful health check with complete service quality.

    Practice Exercise

    Design health checks for your online learning platform's course API.

    Requirements

    1. Create separate startup, liveness, and readiness endpoints.
    2. Define the exact question answered by each endpoint.
    3. Keep liveness independent of the database.
    4. Prevent traffic before configuration loading is complete.
    5. Define database behaviour for course-read operations.
    6. Define degraded behaviour when the recommendation service is unavailable.
    7. Define readiness during connection-pool exhaustion.
    8. Prevent public access to detailed diagnostics.
    9. Add no-cache response headers.
    10. Define probe interval, timeout, and thresholds from test results.
    11. Mark the instance not ready during shutdown.
    12. Drain existing requests before termination.
    13. Test a slow startup.
    14. Test a database outage.
    15. Test an internal application deadlock.
    16. Add an external synthetic course-read test.
    17. Monitor startup duration, readiness changes, and restarts.

    Health-check Design Template

    Check Question Included Conditions Failure Action
    Startup Has initialization completed? Configuration and required local initialization Continue waiting or restart after the approved boundary
    Liveness Can the process make progress? Internal runtime or watchdog state Restart the process
    Readiness Can the instance serve new API requests? Initialization, draining state, and essential serving capacity Remove from active routing
    Detailed diagnostics Which component is degraded? Protected dependency and runtime information Support investigation and alerts
    Synthetic check Can a controlled user journey succeed externally? DNS, TLS, gateway, authentication, routing, and API response Raise an alert or perform approved failover

    Frequently Asked Questions

    1

    What is a health check?

    A health check is an automated probe used to determine whether a process, instance, service, or dependency is in a state suitable for a defined operational action.

    2

    What is a liveness check?

    A liveness check determines whether the application process can continue running and making progress.

    3

    What happens when liveness fails?

    The hosting or orchestration platform commonly restarts the failed process or container according to its configured policy.

    4

    What is a readiness check?

    A readiness check determines whether an instance can safely receive new production traffic.

    5

    What happens when readiness fails?

    The instance is normally removed from active routing without necessarily restarting the process.

    6

    What is a startup check?

    A startup check determines whether application initialization has completed before ordinary liveness and readiness decisions begin.

    7

    Should liveness check the database?

    Generally no. A database outage does not normally mean that restarting the application process will correct the problem.

    8

    Should readiness check the database?

    It depends on whether the database is required for the traffic handled by that instance and whether removing all instances would improve or worsen the failure.

    9

    What is the difference between an L4 and L7 health check?

    An L4 health check normally verifies transport connectivity. An L7 health check sends an application-level request and evaluates the response.

    10

    Should health endpoints return detailed diagnostics?

    Infrastructure-facing endpoints should normally return a minimal verdict. Detailed diagnostics should be protected from unauthorized access.

    11

    Can a health check cause an outage?

    Yes. An overly aggressive or dependency-heavy check can restart healthy instances, remove every backend, amplify dependency traffic, or create repeated health flapping.

    12

    Does a healthy response prove that the service is performing well?

    No. Health indicates suitability for a specific operational decision. Latency, throughput, errors, resource utilization, and user journeys still require monitoring.

    Key Takeaway

    Health checks are operational control signals, not general diagnostic pages. A startup check protects initialization, a liveness check determines whether a process should restart, and a readiness check determines whether an instance should receive new traffic. Keep liveness local and fail it only when a restart can help. Design readiness around the service's actual ability to serve traffic, while considering whether shared dependency failures could remove every backend. Keep frequent checks fast, bounded, and uncached, use thresholds to prevent flapping, and protect detailed diagnostics. Coordinate readiness with deployments and graceful shutdown so new instances receive traffic only after initialization and old instances drain before termination. Finally, combine internal probes with metrics, tracing, alerts, and external synthetic checks to verify the complete user-visible service.