availability and reliability
Availability and Reliability in System Design
Learn how availability measures whether a system is accessible when needed, while reliability measures whether the system consistently performs its intended function correctly over time.
Introduction
Availability and reliability are fundamental quality attributes in system design. They are closely related, but they describe different aspects of a system's behaviour.
- Availability describes whether the system is operational and accessible when users need it.
- Reliability describes whether the system performs its intended function correctly and consistently over a stated period.
A service can be available but unreliable. For example, an API may respond to every request while returning incorrect or incomplete results.
A service can also be reliable when operating but have poor availability. For example, a batch system may always produce correct results but remain inaccessible during long maintenance periods.
Core distinction: Availability asks whether the system can be used. Reliability asks whether the system can be trusted to perform the required operation correctly.
In the System Design curriculum, availability and reliability appear as Topic 1.7 under System Design Process and Estimation, following latency versus throughput and preceding QPS, storage and bandwidth estimation.
Prerequisites
| # | Prerequisite | Why It Is Needed |
|---|---|---|
| 1 | Functional requirements | Reliability is evaluated against the system's intended functions. |
| 2 | Non-functional requirements | Availability and reliability targets are quality requirements. |
| 3 | Request flow | Every required dependency can affect successful request completion. |
| 4 | High-level diagrams | Components and dependencies help identify failure points. |
| 5 | Latency and throughput | A service can technically respond while missing its useful performance target. |
| 6 | Basic probability and percentages | Availability and failure measurements commonly use ratios and probabilities. |
What Is Availability?
Availability is the proportion of required operating time during which a system is capable of providing its intended service.
Availability answers:
A common availability formula is:
\[ Availability = \frac{Uptime}{Uptime + Downtime} \times 100 \]
Where:
- Uptime is the measured time during which the service satisfies the defined availability condition.
- Downtime is the measured time during which the service does not satisfy that condition.
Example Availability Calculation
Suppose a service is expected to operate for 720 hours during a measurement period and experiences 1.5 hours of qualifying downtime.
\[ Uptime = 720 - 1.5 = 718.5\ hours \]
\[ Availability = \frac{718.5}{720} \times 100 = 99.7917\% \]
The percentage is meaningful only when the measurement window, service boundary, qualifying requests and downtime definition are documented.
Define What “Available” Means
A network connection or successful process health check does not automatically prove that the service is available to users.
An availability definition should identify:
- The service or operation being measured
- The users, regions or clients included
- The measurement boundary
- The expected operating schedule
- The acceptable response codes
- The maximum acceptable latency
- Whether degraded responses count as available
- Whether planned maintenance is included
- The minimum request volume needed for evaluation
The application must be available 99.9% of the time.
During:
<defined measurement period and operating schedule>
The operation:
<specific user-visible request>
Shall satisfy:
<defined successful response and latency conditions>
For:
<approved percentage of eligible requests or time>
Measured from:
<defined user or service boundary>
Excluding:
<explicitly approved exclusions>
What Is Reliability?
Reliability is the ability or probability of a system performing its required function correctly, under stated conditions, for a specified period without failure.
Reliability answers:
Reliability includes concerns such as:
- Correct results
- Valid state transitions
- No unintended duplicate business effects
- No lost accepted work
- No corrupted stored data
- Predictable behaviour during retries
- Consistent enforcement of business rules
- Controlled behaviour during component failure
Reliability Example
Payment request received
|
v
Payment processed exactly once
|
v
Order state updated correctly
|
v
Confirmation reflects the actual result
A payment API that remains reachable but sometimes charges the customer twice is available from a connectivity perspective but unreliable from a business perspective.
Availability vs Reliability
| Comparison Area | Availability | Reliability |
|---|---|---|
| Primary question | Can the service be used when required? | Does the service perform the required function correctly? |
| Primary focus | Accessibility and readiness | Correctness and dependable operation |
| Common measurement | Available time or successful-request proportion | Failure rate, success rate or failure-free operating time |
| Improved by | Redundancy, failover, recovery and fault isolation | Correctness, testing, validation, durability and fault prevention |
| Example failure | The API cannot accept requests. | The API accepts a request but produces an incorrect result. |
| User concern | Can the task be attempted? | Can the result be trusted? |
Possible System States
| Availability | Reliability | Example |
|---|---|---|
| High | High | The service remains accessible and consistently produces correct results. |
| High | Low | The service responds but sometimes loses, duplicates or corrupts work. |
| Low | High when operating | The system produces correct results but has frequent or lengthy outages. |
| Low | Low | The system is frequently inaccessible and produces incorrect outcomes. |
MTBF, MTTF and MTTR
Several metrics are commonly used when discussing reliability and availability.
| Metric | Meaning | Interpretation |
|---|---|---|
| MTBF | Mean Time Between Failures | Average operating time between repairable failures |
| MTTF | Mean Time To Failure | Average time before a non-repairable component fails |
| MTTR | Mean Time To Repair or Restore | Average time required to restore service after failure |
| Failure rate | Failures observed over a defined operating quantity | How frequently failures occur |
| Success rate | Successful operations divided by eligible operations | How consistently operations complete successfully |
A common steady-state availability approximation is:
\[ Availability = \frac{MTBF}{MTBF + MTTR} \]
Example
If the average time between failures is 500 hours and the average recovery time is 2 hours:
\[ Availability = \frac{500}{500 + 2} = 0.996016 \]
\[ Availability \approx 99.6016\% \]
This equation shows two broad improvement strategies:
- Increase the time between failures
- Reduce the time required to recover
Availability Targets and Downtime Budgets
An availability target can be converted into a maximum downtime budget for a defined measurement period.
\[ Allowed\ Downtime = Measurement\ Period \times (1 - Availability) \]
| Availability Target | Approximate Maximum Downtime per 365-Day Year |
|---|---|
| 99% | 3 days, 15 hours and 36 minutes |
| 99.9% | 8 hours, 45 minutes and 36 seconds |
| 99.99% | 52 minutes and 34 seconds |
| 99.999% | 5 minutes and 15 seconds |
These values are mathematical approximations. An actual service agreement can use a different measurement period, rounding method, request-based calculation and exclusion policy.
Design warning: Each additional nine can require significant engineering and operational investment. Select targets from business impact and user needs rather than choosing the highest number automatically.
SLA, SLO and SLI
| Term | Meaning | Example |
|---|---|---|
| SLI | Service Level Indicator, a measured signal | Proportion of eligible redirect requests completed successfully |
| SLO | Service Level Objective, a target for an SLI | The approved successful-redirect target over a defined period |
| SLA | Service Level Agreement, a service commitment that can include consequences | A contractual availability commitment |
Architecture should be designed against defined service objectives, while the exact legal or commercial meaning of an SLA belongs to the approved agreement.
Request-based Availability
Availability can also be measured using eligible requests rather than only elapsed time.
\[ Request\ Availability = \frac{Successful\ Eligible\ Requests} {Total\ Eligible\ Requests} \times 100 \]
The definition of a successful eligible request must specify:
- Included endpoints
- Accepted response codes
- Latency threshold
- Client-caused errors
- Rate-limited requests
- Dependency-related failures
- Maintenance exclusions
- Regions and clients included
Serial Dependency Availability
If a request requires several independent components and the request fails when any required component is unavailable, overall availability decreases as dependencies are added.
For independent required components:
\[ A_{system} = A_1 \times A_2 \times \cdots \times A_n \]
Example
If a request always requires three independent components, each with availability of 99.9%:
\[ A_{system} = 0.999 \times 0.999 \times 0.999 \]
\[ A_{system} \approx 0.997003 = 99.7003\% \]
This simplified model demonstrates why long synchronous dependency chains can reduce end-to-end availability.
Redundancy
Redundancy provides more than one component capable of performing a required function.
+--> [Service Instance A]
[Load Balancer] ---+
+--> [Service Instance B]
|
+--> [Service Instance C]
Redundancy can reduce the effect of a single component failure when:
- Traffic is routed away from unhealthy instances
- The instances do not share the same failure condition
- Required state is available to replacement instances
- Capacity remains sufficient after an instance fails
- Detection and failover occur within the required recovery target
Important: Multiple copies do not guarantee high availability when every copy depends on the same network, database, configuration, credential or deployment fault.
Failover
Failover transfers work from an unavailable or unhealthy component to an alternative.
Normal:
Traffic -> Primary Component
Failure detected:
Primary Component -> Unhealthy
Failover:
Traffic -> Secondary Component
A failover design should define:
- How failure is detected
- How false failure detection is controlled
- Who initiates failover
- How traffic or leadership moves
- Whether data is current enough
- How long failover takes
- Whether in-progress work is lost or repeated
- How recovery to the normal state occurs
- How failover is tested
Fault Tolerance
Fault tolerance is the ability of a system to continue providing an acceptable service despite one or more component failures.
Fault-tolerance techniques include:
- Redundant instances
- Replicated data
- Automatic failover
- Retry with bounded backoff
- Idempotent processing
- Queue-based buffering
- Circuit breakers
- Bulkheads
- Load shedding
- Graceful degradation
Fault tolerance does not mean hiding every failure. The system can reject selected work or disable optional features to protect critical functions.
Fault Isolation and Bulkheads
Fault isolation prevents a failure in one component, tenant or workload from consuming all shared resources.
Unsafe resource sharing:
Redirect Traffic ----+
+--> Shared Worker Pool
Analytics Traffic ---+
|
+--> Shared Database Connections
Analytics overload can affect redirects.
Isolated resources:
Redirect Traffic --> Redirect Pool --> Critical Read Store
Analytics Traffic --> Analytics Pool --> Analytics Store
Isolation boundaries can be created through:
- Separate worker pools
- Bounded queues
- Independent connection pools
- Resource quotas
- Separate failure domains
- Tenant limits
- Independent deployments
Graceful Degradation
Graceful degradation allows critical capabilities to remain available when an optional or lower-priority component fails.
URL Shortener Example
Redirect request
|
v
Redirect mapping found
|
+--> Return redirect
|
+--> Attempt analytics publication
|
+--> Success: analytics continues
|
+--> Failure: follow approved analytics-failure policy
If immediate analytics recording is not part of the redirect's critical correctness requirement, analytics failure may not need to block the redirect.
Possible degraded modes include:
- Serving cached read-only data
- Disabling optional recommendations
- Delaying analytics processing
- Rejecting large background jobs
- Serving a reduced response
- Temporarily disabling noncritical administration features
Reliability and Data Correctness
Reliability is not limited to keeping processes running. It also requires correct data handling.
Important data-reliability concerns include:
- Atomicity of related state changes
- Valid transaction boundaries
- Duplicate-request protection
- Durable storage of accepted work
- Detection of corrupted data
- Backup and restoration
- Correct cache invalidation
- Ordering of dependent operations
- Reconciliation after uncertain outcomes
Uncertain Write Example
Client sends create request
|
v
Database commits record
|
v
Response is lost
|
v
Client observes timeout
|
v
Client retries
|
+--> Without protection: duplicate record
|
+--> With idempotency: return original result
An idempotency mechanism can protect reliability when a caller cannot know whether a timed-out operation completed.
Retries and Reliability
A retry can improve recovery from a temporary failure, but uncontrolled retries can reduce availability by increasing load.
A safe retry policy should specify:
- The retry owner
- Retryable failures
- Maximum attempts
- Backoff strategy
- Randomized delay where appropriate
- Remaining request deadline
- Idempotency requirements
- Overload protection
- Observability and alerting
Client retries
Gateway retries
Application retries
Database library retries
Worker retries
Retry owner:
Application service
Conditions:
Only approved temporary failures
Attempts:
Bounded by policy
Delay:
Backoff with controlled randomness
Safety:
Idempotent operation or idempotency key
Deadline:
Retry only while sufficient time remains
Complete URL Shortener Example
Redirect Availability
Visitor
|
v
Edge / Load Balancer
|
v
Redirect Service Instances
|
+--> Redirect Cache
|
+--> Link Mapping Store
|
v
Redirect Response
Availability considerations include:
- Multiple healthy redirect-service instances
- Health-aware traffic routing
- Sufficient capacity after an instance failure
- Defined cache-failure behaviour
- Replicated or recoverable mapping storage
- Timeouts for storage access
- Controlled behaviour when no mapping can be retrieved
Redirect Reliability
Reliability considerations include:
- A short code resolves to the correct destination.
- An expired link does not remain active beyond the approved visibility delay.
- A deactivated link follows the documented cache-consistency rule.
- An unknown code produces the correct not-found result.
- A mapping is not corrupted during creation or migration.
- Analytics processing does not silently alter redirect correctness.
- Retries do not create duplicate link mappings.
Degraded Redirect Flow
Redirect request
|
v
Cache lookup
|
+--> Hit
| |
| v
| Validate cached mapping
| |
| v
| Redirect response
|
+--> Miss
|
v
Link Store
|
+--> Available: retrieve mapping
|
+--> Unavailable:
|
+--> Approved stale-data path, if permitted
|
+--> Controlled failure
Serving stale data must be an explicit policy decision. It is not appropriate when stale state can produce an unsafe or incorrect outcome.
Availability and Consistency Trade-offs
During network or storage failures, a distributed system may be unable to simultaneously provide every consistency and availability expectation.
For each important operation, ask:
- Can the system serve stale data?
- Can the operation be rejected temporarily?
- Can a write be accepted for later processing?
- Must the caller observe the latest committed value?
- Can conflicting updates occur?
- How will conflicts be resolved?
- What must the user observe during partial failure?
The answer can differ by workflow. A content feed may tolerate temporary staleness, while balance-changing operations may require stricter correctness.
Single Points of Failure
A single point of failure is a component whose failure can make the required service unavailable because no acceptable alternative path exists.
[Client]
|
v
[Single Gateway]
|
v
[Application Instances]
|
v
[(Database)]
Possible single points:
- One gateway instance
- One database instance
- One network path
- One regional dependency
- One credential or configuration source
Reviewing single points of failure requires examining not only server count, but also shared dependencies and correlated failures.
Failure Domains
A failure domain is a set of components that can fail together because they share infrastructure or dependencies.
Examples include:
- One process
- One virtual machine
- One physical host
- One network segment
- One availability zone
- One region
- One identity or configuration system
- One deployment or software version
Replicas should be placed and operated so that one failure does not disable every replica simultaneously.
Planned Maintenance
Availability design must account for maintenance, deployment, schema changes and infrastructure replacement.
Maintenance-friendly designs can include:
- Rolling deployments
- Backward-compatible API and schema changes
- Health-aware traffic removal
- Capacity for one or more instances to be unavailable
- Maintenance windows where required
- Tested rollback procedures
- Version compatibility during staged releases
Recovery Objectives
Availability and reliability planning commonly includes recovery time and recovery point objectives.
| Objective | Question |
|---|---|
| Recovery Time Objective | How quickly must the service or data capability be restored? |
| Recovery Point Objective | How much data loss, measured in time, is acceptable after recovery? |
Backup creation alone does not prove recoverability. Restoration must be tested against the documented recovery objectives.
Monitoring Availability and Reliability
| Signal | What It Helps Detect |
|---|---|
| Successful-request ratio | User-visible request failures |
| Latency percentiles | Requests that are technically successful but too slow |
| Error rate by operation | Failing workflows and components |
| Timeout rate | Dependency delay and capacity pressure |
| Queue depth and oldest-message age | Asynchronous backlog and delayed completion |
| Replication lag | Stale reads and recovery exposure |
| Data-integrity checks | Missing, duplicated or corrupted state |
| Failover and recovery events | Frequency and effectiveness of resilience mechanisms |
| Resource saturation | Capacity conditions that can trigger availability loss |
Reliability and Availability Testing
| Test Type | Purpose |
|---|---|
| Functional test | Verify correct behaviour under normal conditions |
| Negative test | Verify safe rejection of invalid operations |
| Load test | Verify availability and correctness under expected workload |
| Stress test | Identify behaviour beyond sustainable capacity |
| Failover test | Verify detection, traffic movement and recovery |
| Dependency-failure test | Verify timeout, retry, fallback and degraded behaviour |
| Backup restoration test | Verify data can be restored within approved objectives |
| Data-integrity test | Detect missing, duplicated or corrupted records |
| Soak test | Verify stable operation over an extended workload |
| Controlled fault experiment | Verify resilience assumptions under an approved failure scenario |
Failure Test Template
Test ID:
FT-<number>
Related requirement:
<Availability or reliability requirement>
Failure introduced:
<Component, dependency or network condition>
Preconditions:
<Required system state and workload>
Expected detection:
<Health check, alert or timeout>
Expected system response:
<Failover, fallback, rejection or degraded mode>
Expected user behaviour:
<What the user observes>
Expected data behaviour:
<No loss, bounded loss, no duplicate or reconciliation rule>
Recovery:
<How normal operation is restored>
Measurements:
- Detection time
- Recovery time
- Error rate
- Latency
- Lost or duplicated operations
- Remaining capacity
Result:
Passed / Failed / Partially passed
Common Mistakes
Treating Availability and Reliability as Synonyms
Availability concerns accessibility. Reliability concerns correct and dependable operation.
Measuring Only Process Uptime
A running process does not prove that users can complete the required operation successfully.
Counting Incorrect Responses as Available
Define success in user-visible terms, including response correctness and useful latency.
Adding Replicas in the Same Failure Domain
Replicas sharing the same host, network, configuration or dependency can fail together.
Ignoring Capacity After Failure
Remaining instances must have enough capacity to serve the required workload after failover.
Retrying Every Failure
Uncontrolled retries can increase overload and reduce availability.
Failing Open When Correctness Requires Rejection
Some operations should reject requests rather than return stale or unverified results.
Making Optional Dependencies Critical
Analytics or recommendation failure should not block a critical workflow unless synchronous completion is explicitly required.
Having Backups Without Restoration Tests
Recovery confidence requires evidence that backups can be restored within the defined objectives.
Choosing a High Availability Target Without Justification
Higher targets increase engineering and operating requirements. Select them according to user and business impact.
Ignoring Planned Changes
Deployments, schema changes and maintenance can cause outages if compatibility and rollback are not designed.
Failing to Test Degraded Modes
A fallback that exists only in documentation is not a verified resilience capability.
Design Review Checklist
Availability and Reliability Review
- The critical user journeys are identified.
- Availability is defined at a user-visible boundary.
- Reliability defines correct business outcomes.
- The measurement period is documented.
- Latency thresholds are included where required.
- Planned-maintenance treatment is documented.
- Single points of failure are identified.
- Correlated failure domains are identified.
- Redundant components have independent failure characteristics.
- Capacity after failure is validated.
- Health checks represent useful service behaviour.
- Failover rules are documented and tested.
- Timeout and retry ownership is explicit.
- Non-idempotent operations have duplicate protection.
- Optional functionality is separated from critical paths.
- Graceful degradation is defined where safe.
- Data integrity and durability are monitored.
- Backups and restoration are tested.
- Recovery objectives are documented.
- Availability and reliability objectives are monitored continuously.
Practice Exercise
Analyze availability and reliability for a notification platform that accepts email, SMS and mobile push requests.
System Flow
Client
|
v
Notification API
|
v
Notification Store
|
v
Delivery Queue
|
v
Channel Worker
|
v
External Provider
|
v
Delivery Status Store
Tasks
- Define availability for notification submission.
- Define reliability for accepted notifications.
- Identify every single point of failure.
- Define behaviour when the delivery provider is unavailable.
- Define retryable and permanent failures.
- Explain how duplicate delivery is prevented or handled.
- Define the required status when delivery completion is uncertain.
- Identify metrics for queue backlog and delivery failure.
- Create one failover test.
- Create one recovery test.
Model Requirements
Availability requirement:
During the defined operating period, the notification-submission API shall
satisfy the approved successful-request and latency objectives for eligible
requests measured at the public API boundary.
Reliability requirement:
After the system confirms durable acceptance, the notification shall reach a
documented terminal state such as delivered, permanently failed or cancelled.
Retries shall follow the approved duplicate-handling and delivery policy.
Degraded-mode requirement:
A temporary provider outage shall not cause an accepted notification to be
silently lost. The notification shall remain pending, retry according to
policy or move to the documented failure-handling path.
Frequently Asked Questions
What is availability?
Availability measures whether a system or required operation is accessible and usable when needed.
What is reliability?
Reliability measures whether a system performs its required function correctly and consistently under stated conditions over time.
Can a system be available but unreliable?
Yes. An online service can respond to requests while producing incorrect, incomplete or duplicated results.
Can a system be reliable but unavailable?
Yes. A system may produce correct results whenever it operates but remain inaccessible for an unacceptable amount of time.
Does redundancy guarantee availability?
No. Redundancy helps only when alternatives do not fail together, traffic can reach them and sufficient capacity and valid state remain available.
What is the difference between reliability and durability?
Reliability covers correct and dependable system operation. Durability specifically concerns preserving accepted data despite failures.
What is graceful degradation?
Graceful degradation preserves critical functionality while reducing or disabling noncritical functionality during failure or overload.
What is a single point of failure?
It is a component or dependency whose failure makes the required service unavailable because no acceptable alternative path exists.
Should every service target five nines of availability?
No. Availability targets should reflect user needs, business impact, cost, dependencies and operational capability.
What comes after availability and reliability?
The next step is to estimate QPS, storage and bandwidth using documented workload assumptions before sizing major system components.
Key Takeaway
Availability measures whether users can access and use a required service, while reliability measures whether the system performs that service correctly and consistently. Define both attributes at user-visible boundaries, identify failure domains and single points of failure, design redundancy and failover carefully, isolate optional work, protect data correctness and test recovery rather than assuming it works. A dependable architecture must remain accessible enough for its purpose without sacrificing the correctness of the results it provides.