Table of Contents

    durability

    STORAGE, FILES, OBJECTS & SEARCH BASICS

    Durability

    Learn how storage systems preserve acknowledged data across hardware failures, process crashes, corruption, infrastructure outages, accidental deletion, and disasters through replication, erasure coding, checksums, transaction logs, snapshots, backups, versioning, and tested recovery procedures.

    Introduction

    Storing data successfully is not enough. A reliable storage system must preserve that data after the client receives confirmation that the write has completed.

    Durability describes the system's ability to keep acknowledged data intact and recoverable according to its documented failure model.

    Durability concerns include:

    • Process crashes
    • Operating-system failures
    • Storage-device failures
    • Power interruptions
    • Network partitions
    • Data corruption
    • Availability-zone failures
    • Regional disasters
    • Accidental deletion
    • Malicious modification
    • Software defects

    Core idea: Durability is meaningful only when the system defines which failures it protects against, when a write is considered successful, how many copies or recovery structures exist, and how the data can be verified and restored.

    In your System Design curriculum, Durability is Topic 6.2 under Storage, Files, Objects & Search Basics. It follows block, file, and object storage and precedes checksums, metadata, uploads and multipart transfer, inverted indexes, and full-text search.

    Prerequisites

    # Prerequisite Why It Is Needed
    1 Block, file, and object storage Durability mechanisms apply differently to volumes, files, and objects.
    2 Basic failure concepts Durability guarantees are defined against specific failure scenarios.
    3 Replication Multiple copies can protect against the loss of one storage component.
    4 Transactions Committed database changes require durable logging and recovery.
    5 Checksums Integrity verification helps detect corrupted stored or transferred data.
    6 Backup and recovery Historical recovery protects against failures that replication alone cannot address.

    What Is Durability?

    Durability is the guarantee that, after a storage system acknowledges a successful write, the stored data remains protected and recoverable under the failures covered by that system's durability contract.

    Client sends write
          |
          v
    Storage system receives data
          |
          v
    Required persistence or redundancy completes
          |
          v
    System acknowledges success
          |
          v
    Covered failure occurs
          |
          v
    Acknowledged data remains present
    or can be recovered

    Durability is not an unlimited promise. A single-device system, replicated storage service, versioned object store, and geographically distributed backup system protect against different failure scopes.

    Durability Contract
    acknowledged write → defined failure occurs → data remains intact → recovery is possible

    Durability vs Availability

    Durability and availability answer different questions.

    Property Primary Question Example
    Durability Has the acknowledged data remained intact? The file still exists after a storage-device failure
    Availability Can the data be accessed when requested? The file can be downloaded immediately
    Durable but temporarily unavailable:
    
    The data remains safely stored,
    but a network outage prevents access.
    
    
    Available but insufficiently durable:
    
    The data is currently readable,
    but only one vulnerable copy exists.

    Design rule: High availability does not automatically prove high durability, and durable data can still become temporarily unavailable.

    Durability vs Consistency

    Property Concern
    Durability Whether acknowledged data survives covered failures
    Consistency Which valid value or state readers observe
    Integrity Whether stored content remains correct and uncorrupted

    A replica can durably retain an older value while another replica stores a newer value. Durability alone does not define which version a reader will observe.

    Define the Failure Model

    A durability guarantee should identify the failures it is designed to tolerate.

    Failure Example Protection
    Application process crash Write through a durable storage service before acknowledgment
    Operating-system crash Persistent journal or transaction log
    Single drive failure Replication or erasure coding across devices
    Server failure Copies on independent servers
    Availability-zone failure Copies or encoded fragments across zones
    Regional disaster Geographically separate replica or backup
    Accidental deletion Versioning, snapshots, or independent backup
    Silent corruption Checksums, verification, and repair from redundant data
    Malicious encryption or deletion Protected immutable or isolated recovery copies

    Write Acknowledgment

    The moment at which the system confirms a write is central to its durability guarantee.

    Weak acknowledgment:
    
    1. Receive write.
    2. Keep data only in volatile memory.
    3. Return success.
    4. Process loses power.
    5. Acknowledged data can be lost.
    
    
    Stronger acknowledgment:
    
    1. Receive write.
    2. Persist required recovery information.
    3. Create required redundant representation.
    4. Return success.

    Stronger acknowledgment can require additional storage operations or network round trips, increasing write latency.

    Volatile vs Persistent Storage

    Storage General Property
    Volatile memory Content can be lost when power or the process is lost
    Persistent storage Designed to retain content across process or power interruption
    Replicated persistent storage Maintains redundant data representations across failure domains
    Backup storage Maintains recoverable historical copies separately from active data

    Caches improve performance but should not be treated as the only durable copy unless the cache is explicitly designed and configured as a durable storage system.

    Write-ahead Logging

    A database can use write-ahead logging so that recovery information is made durable before modified data pages are written to their final locations.

    Transaction modifies data
          |
          v
    Create transaction-log records
          |
          v
    Persist required log records
          |
          v
    Acknowledge commit
          |
          v
    Write modified data pages later

    After a crash, the database can use its transaction log to restore committed operations and remove incomplete work according to its recovery algorithm.

    Conceptual Database Transaction

    BEGIN;
    
    INSERT INTO enrollments
    (
        enrollment_id,
        learner_id,
        course_id,
        enrollment_status
    )
    VALUES
    (
        :enrollment_id,
        :learner_id,
        :course_id,
        'active'
    );
    
    COMMIT;

    The database durability configuration determines what persistence work must complete before the commit is acknowledged.

    Replication

    Data replication maintains multiple copies of stored data, called replicas, across selected storage components or locations.

    Client write
        |
        v
    Primary storage
        |
        +-- Replica A
        |
        +-- Replica B
        |
        +-- Replica C

    Replication can help protect against drive, server, rack, zone, or regional failure depending on where replicas are placed.

    Synchronous Replication

    With synchronous replication, the write acknowledgment waits for the required replica confirmations.

    Client sends write
          |
          v
    Primary stores write
          |
          v
    Required replica stores write
          |
          v
    Replica acknowledges
          |
          v
    Primary returns success

    Potential Benefit

    The acknowledged write is already represented on the required participating replicas.

    Consideration

    Write latency and availability can depend on the slowest required replica or network path.

    Asynchronous Replication

    With asynchronous replication, the primary can acknowledge a write before a remote replica receives it.

    Client sends write
          |
          v
    Primary stores write
          |
          v
    Primary returns success
          |
          v
    Replication continues asynchronously

    If the primary is lost before replication completes, recently acknowledged data can be absent from the remote replica, depending on the service guarantee.

    Replication Comparison

    Area Synchronous Asynchronous
    Acknowledgment Waits for required replica persistence Can occur before remote replication completes
    Write latency Includes required replication path Can avoid waiting for remote copy
    Recent-data exposure Lower for covered replica failures Replication lag can expose recent writes
    Failure handling Required replica availability can affect writes Primary can continue while replica catches up

    Quorum-based Persistence

    Distributed storage can acknowledge a write after a required subset of replicas confirms it.

    Replication factor:
    
    N
    
    
    Required write acknowledgments:
    
    W
    
    
    Required read responses:
    
    R

    A simplified condition often discussed for overlapping read and write quorums is:

    \[ R + W > N \]

    This equation alone does not fully describe the consistency or durability of a real system. Replica failures, conflict resolution, repair, timing, and implementation details also matter.

    Erasure Coding

    Erasure coding divides data into fragments and generates additional parity fragments. The original content can be reconstructed when enough fragments remain available.

    Original data
          |
          v
    Split into data fragments
          |
          v
    Generate parity fragments
          |
          v
    Distribute fragments across devices
    
    
    After selected fragment loss:
    
    Remaining fragments
          |
          v
    Reconstruct original content

    Erasure coding can provide storage-efficient redundancy but requires encoding, distribution, reconstruction, and repair work.

    Replication vs Erasure Coding

    Area Replication Erasure Coding
    Representation Complete copies of data Data and parity fragments
    Recovery Read from another complete replica Reconstruct from available fragments
    Storage overhead Depends on number of full copies Depends on data and parity configuration
    Repair work Copy from a healthy replica Read and reconstruct required fragments
    Typical decision factors Latency, simplicity, and workload Capacity efficiency and large-scale durability

    Checksums and Integrity

    Durability must protect not only against missing data but also against undetected corruption.

    Write data
        |
        v
    Calculate checksum
        |
        v
    Store data and checksum
        |
        v
    Read data later
        |
        v
    Recalculate checksum
        |
        +-- Match:
        |      content passed verification
        |
        +-- Mismatch:
               corruption detected

    When redundant data exists, the system can use a valid copy or fragment set to repair corrupted content.

    File-checksum Example

    sha256sum course-video.mp4

    Checksums and integrity verification are covered in detail in the next topic after durability.

    Scrubbing

    Data scrubbing is the periodic process of reading stored data, verifying its integrity, and repairing detected corruption when valid redundant data is available.

    Periodic verification job
            |
            v
    Read stored fragment or object
            |
            v
    Verify checksum
            |
            +-- Valid:
            |      continue
            |
            +-- Invalid:
                   locate healthy copy
                   repair corrupted data
                   record event

    Without periodic verification, silent corruption can remain undiscovered until recovery is attempted.

    Snapshots

    A snapshot captures storage state at a specific point in time.

    Live volume
        |
        +-- Snapshot at Time A
        |
        +-- Snapshot at Time B
        |
        +-- Snapshot at Time C

    Snapshots can support:

    • Recovery from accidental modification
    • Volume cloning
    • Testing against a captured state
    • Backup workflows
    • Point-in-time recovery strategies

    A snapshot stored in the same failure domain can share risks with the source storage. Its independence and recovery value must be verified.

    Object Versioning

    Object versioning retains previous generations when an object is overwritten or deleted according to the configured policy.

    Object key:
    
    courses/42/lesson-7/notes.pdf
    
    
    Version 1:
    Original content
    
    
    Version 2:
    Corrected content
    
    
    Version 3:
    Latest content

    Versioning can help recover from accidental overwrite or deletion, but it increases storage consumption and requires lifecycle and retention rules.

    Backups

    A backup is a timestamped copy of data that can be used to restore lost or damaged information.

    Production storage
            |
            v
    Backup process
            |
            v
    Timestamped recovery copy
            |
            v
    Protected backup location

    Backup design should define:

    • What data is included
    • How often copies are created
    • Where copies are stored
    • How long copies are retained
    • How copies are protected
    • How restoration is performed
    • How restoration is tested

    Replication vs Backup

    Area Replication Backup
    Primary purpose Maintain additional current copies Maintain recoverable historical copies
    Accidental deletion Deletion can propagate to replicas An earlier backup can retain the deleted data
    Corruption Corruption can replicate if not detected A clean earlier copy can support recovery
    Availability Can support failover to another copy Usually requires a restore operation
    Historical recovery Not necessarily provided A primary purpose of backup retention

    Recovery rule: Replication improves redundancy, but it does not replace protected backups. Accidental deletion, corruption, or a faulty application update can affect multiple replicas.

    Backup Isolation and Immutability

    A recovery copy should be protected from the same identities, errors, or failures that can damage production data.

    Protection can include:

    • Separate security boundaries
    • Restricted deletion permissions
    • Retention locks where required
    • Immutable backup copies
    • Independent credentials
    • Geographically separate storage
    • Audited recovery operations

    Recovery Point Objective

    Recovery Point Objective, commonly abbreviated as RPO, describes the acceptable amount of data loss measured in time.

    Last recoverable point:
    
    10:00
    
    
    Failure:
    
    10:15
    
    
    Potential lost interval:
    
    15 minutes

    A smaller RPO requires more frequent or more continuous protection.

    Recovery Time Objective

    Recovery Time Objective, commonly abbreviated as RTO, describes the target time for restoring an acceptable level of service after a disruption.

    Failure detected
          |
          v
    Recovery begins
          |
          v
    Data restored or failover completed
          |
          v
    Service becomes usable

    Durability protects data, while RTO addresses how quickly the service must become operational again.

    RPO vs RTO

    Objective Question
    RPO How much recent data loss is acceptable?
    RTO How quickly must the service or data be restored?

    Failure Domains

    Redundant copies should be placed across failure domains appropriate to the required durability target.

    Device
      |
      v
    Server
      |
      v
    Rack
      |
      v
    Availability zone
      |
      v
    Region
      |
      v
    Independent backup location

    Two copies on one physical device do not protect against that device's failure. Several copies in one building do not protect against the complete loss of that building.

    Regional and Geographic Redundancy

    Geographic redundancy stores or replicates data in locations separated from the primary storage location.

    Design questions include:

    • Is replication synchronous or asynchronous?
    • What recent-data exposure remains?
    • Can applications read from the secondary location?
    • How is failover initiated?
    • How is failback performed?
    • Are encryption keys available in the recovery location?
    • Are regulatory location requirements satisfied?

    Database Durability

    Database durability normally combines transaction logs, checkpoints, persistent data pages, recovery processing, and the underlying storage platform.

    Application transaction
            |
            v
    Database transaction log
            |
            v
    Durable storage acknowledgment
            |
            v
    COMMIT returned
            |
            v
    Checkpoint and page persistence
            |
            v
    Recovery after failure

    Database durability settings can trade write latency for a weaker or stronger acknowledgment guarantee. Review platform-specific configuration before changing the default.

    File-storage Durability

    File-storage durability can combine protected underlying devices, file-system journaling, redundant servers, snapshots, checksums, and backups.

    Application writes file
          |
          v
    File-system operation
          |
          v
    File metadata and data persisted
          |
          v
    Storage redundancy applied
          |
          v
    Backup or snapshot retained

    A successful application write does not automatically prove that all buffers have reached durable media. File APIs, flush operations, mount settings, and storage guarantees must be understood.

    Object-storage Durability

    An object-storage system can protect data by storing redundant complete copies or encoded fragments across several devices or locations.

    Object upload
          |
          v
    Validate request and data
          |
          v
    Store redundant representation
          |
          v
    Persist object metadata
          |
          v
    Acknowledge successful upload
          |
          v
    Periodically verify integrity

    Object versioning, retention, lifecycle policies, and independent backup can protect against additional risks such as accidental overwrite or deletion.

    Database and Object-storage Workflow

    An object upload and relational database update normally cannot share one ordinary local transaction.

    Upload object succeeds
            |
            v
    Database metadata insert fails
            |
            v
    Durable object exists
    without active business record

    A resilient workflow can use:

    • Pending metadata status
    • Generated object identifiers
    • Upload verification
    • Idempotent finalization
    • Abandoned-object cleanup
    • Metadata-to-storage reconciliation

    Durable Upload Workflow

    1. Create pending asset metadata.
    
    2. Generate a stable object key.
    
    3. Upload the object.
    
    4. Verify object existence,
       expected size,
       and checksum.
    
    5. Mark the asset active.
    
    6. Retain appropriate versions or backups.
    
    7. Reconcile incomplete workflows.

    Accidental-deletion Recovery

    Object deleted accidentally
            |
            v
    Check object versions
            |
            +--> Previous version exists:
            |       restore selected version
            |
            +--> No retained version:
                    locate backup
                    restore object
                    verify checksum
                    verify metadata reference

    A deletion process should account for versions, snapshots, replicas, backups, caches, search indexes, legal holds, and business metadata.

    Restore Testing

    A backup is useful only when the required data can be restored correctly.

    A restore test should verify:

    • The backup can be located
    • Decryption keys are available
    • The backup can be read
    • The required recovery point exists
    • The restored data passes integrity checks
    • Application references remain valid
    • The restored service can start
    • Access control remains correct
    • Recovery completes within the required objective

    Operational rule: Do not treat a successful backup job as proof of recoverability. Perform controlled restore tests and verify the restored data.

    Durability Metrics

    A durability percentage communicates a statistical design target for data loss over a period. It is not a promise that an individual object can never be lost.

    A simplified annual loss expectation can be expressed as:

    \[ ExpectedLostObjects = StoredObjects \times AnnualLossProbability \]

    Use provider-published definitions when interpreting reported durability. Confirm the measured unit, period, exclusions, service scope, and failure assumptions.

    Durability and Data Retention

    Durable storage preserves data, while retention rules determine how long the data should remain.

    Concern Question
    Durability Can the stored data survive covered failures?
    Retention How long should the stored data or versions be retained?
    Deletion When and how should data be removed?
    Legal hold Must normal expiration or deletion be suspended?
    Archival Can older data move to a lower-access storage tier?

    Keeping data forever is not automatically safer. Unnecessary retention can increase privacy, compliance, operational, and cost risks.

    Encryption-key Durability

    Encrypted durable content becomes unusable if the required decryption keys are permanently lost.

    Durable encrypted backup
            |
            v
    Decryption key unavailable
            |
            v
    Data cannot be restored

    Key-management planning should address:

    • Key backup or protected redundancy
    • Access control
    • Rotation
    • Recovery-location availability
    • Auditability
    • Revocation and destruction

    Threats to Durable Data

    Durability planning should consider:

    • Hardware failure
    • Software defects
    • Operator mistakes
    • Expired credentials
    • Lost encryption keys
    • Ransomware
    • Unauthorized deletion
    • Silent corruption
    • Incorrect lifecycle policies
    • Replication lag
    • Incomplete backups
    • Untested restoration

    Durability Observability

    Useful durability metrics include:

    • Write acknowledgment failures
    • Replication lag
    • Under-replicated object or block count
    • Checksum verification failures
    • Repair backlog
    • Storage-device failures
    • Snapshot success and failure count
    • Backup success and failure count
    • Backup age
    • Restore-test success rate
    • Restore duration
    • Unrecoverable-object count
    • Expired-version count
    • Recovery-point age

    Alerting

    Alert conditions can include:

    • Required replica count is not satisfied
    • Replication lag exceeds the approved threshold
    • Backup did not complete
    • No recent valid recovery point exists
    • Checksum mismatches increase
    • Repair backlog grows
    • Restore test fails
    • Backup or archive storage approaches capacity
    • Retention or immutability policy changes unexpectedly

    Durability-review Workflow

    1. Identify the data and its business importance.
    2. Define the acknowledged-write semantics.
    3. Define the failures the system must tolerate.
    4. Identify storage and replica failure domains.
    5. Define required replication or erasure coding.
    6. Define checksum and integrity verification.
    7. Define snapshot and versioning policies.
    8. Define backup frequency and retention.
    9. Define RPO and RTO.
    10. Protect recovery copies and encryption keys.
    11. Implement repair and reconciliation.
    12. Test failure and restoration procedures.
    13. Monitor replication, backups, integrity, and recovery readiness.

    Common Durability Mistakes

    1

    Acknowledging before Required Persistence

    A failure can remove data that the client was told had been stored successfully.

    2

    Keeping Only One Copy

    One device, server, or location can remain a single point of data loss.

    3

    Placing Every Replica in One Failure Domain

    Several copies cannot protect against a complete failure of their shared location or dependency.

    4

    Assuming Replication Is Backup

    Deletion, corruption, and faulty updates can propagate to replicas.

    5

    Never Verifying Checksums

    Silent corruption can remain undetected until the content is needed.

    6

    Creating Backups without Restore Tests

    A completed backup job does not prove that the data, keys, and application can be recovered.

    7

    Storing Backups with Production Permissions

    The same compromised identity can delete or corrupt both production data and its recovery copies.

    8

    Ignoring Replication Lag

    A secondary copy can be missing recent writes when failover occurs.

    9

    Ignoring Encryption Keys

    Durable encrypted data is not recoverable when its required keys are lost.

    10

    Keeping Every Version Forever

    Unlimited retention increases cost and can violate privacy or deletion requirements.

    11

    Confusing Durability with Availability

    Data can remain intact while being temporarily inaccessible.

    12

    Using Provider Durability Numbers without Understanding Scope

    A published target must be interpreted using its documented period, workload, exclusions, failure model, and service configuration.

    Recommended Test Cases

    Test Expected Evidence
    Process crash before acknowledgment The client does not receive a false successful result
    Process crash after acknowledged commit The committed data remains recoverable
    Single-device failure Data is reconstructed or read from healthy redundancy
    Replica failure The under-replicated state is detected and repaired
    Checksum mismatch Corruption is detected and not returned silently
    Accidental object overwrite The required previous version can be recovered
    Accidental deletion The object or file can be restored according to policy
    Backup failure Monitoring identifies the missing recovery point
    Restore test Data is restored and passes integrity validation
    Lost encryption-key simulation The key-recovery procedure is validated safely
    Regional recovery exercise Documented data and service recovery objectives are measured
    Metadata-object reconciliation Missing and orphaned objects are detected

    Durability Best Practices

    Recommended Practices

    • Define the durability failure model explicitly.
    • Acknowledge writes only after required persistence completes.
    • Use redundancy appropriate to the required failure domains.
    • Choose synchronous and asynchronous replication deliberately.
    • Monitor replication lag and under-replicated data.
    • Use checksums to detect corruption.
    • Periodically verify stored data and repair damaged copies.
    • Use snapshots or versions where point-in-time recovery is required.
    • Maintain protected historical backups.
    • Separate backup security from production access where practical.
    • Define RPO and RTO from business requirements.
    • Protect encryption keys required for recovery.
    • Apply approved retention and deletion policies.
    • Reconcile database metadata with files and objects.
    • Clean orphaned and abandoned storage safely.
    • Test process, device, zone, and recovery failures where approved.
    • Perform recurring restore tests.
    • Verify restored data using checksums and application checks.
    • Monitor backup age, integrity failures, repair backlog, and restore readiness.
    • Document ownership for recovery decisions and procedures.

    Practice Exercise

    Design a durability and recovery plan for your online learning platform.

    Requirements

    1. Classify course videos, documents, certificates, database records, and audit logs.
    2. Define the acknowledged-write point for each data type.
    3. Identify device, server, zone, and regional failure requirements.
    4. Define replication or erasure-coding requirements.
    5. Store and verify checksums for uploaded content.
    6. Define object-versioning rules.
    7. Define snapshot and backup schedules.
    8. Define recovery retention periods.
    9. Set RPO and RTO for critical platform data.
    10. Protect backup credentials and encryption keys.
    11. Detect orphaned objects and metadata records.
    12. Test accidental object deletion.
    13. Test database restoration.
    14. Test recovery of one learner certificate.
    15. Record recovery results and corrective actions.

    Durability-design Template

    Data Failure Model Protection Recovery Objective
    Course database Process, device, zone, and operator failure Transaction log, replication, backup, and restore testing Record approved RPO and RTO
    Lesson videos Device corruption, deletion, and regional disruption Object redundancy, checksum, versioning, and backup Record approved recovery requirements
    Learner certificates Deletion, corruption, and metadata mismatch Object protection, checksum, and regeneration or backup Record approved recovery requirements
    Audit records Unauthorized deletion and storage failure Controlled retention, protected copies, and integrity checks Record approved retention and recovery rules
    Search index Index loss or corruption Rebuild from authoritative source where supported Record rebuild target

    Frequently Asked Questions

    1

    What is storage durability?

    Storage durability is the guarantee that acknowledged data remains intact and recoverable under the failures defined by the storage system.

    2

    Is durability the same as availability?

    No. Durability concerns whether data remains intact, while availability concerns whether the data can be accessed when requested.

    3

    How does replication improve durability?

    Replication maintains multiple copies so data can survive the loss of a covered storage component or location.

    4

    What is synchronous replication?

    Synchronous replication waits for required replica persistence before acknowledging the write.

    5

    What is asynchronous replication?

    Asynchronous replication allows the primary to acknowledge a write before remote replication is complete.

    6

    What is erasure coding?

    Erasure coding stores data and parity fragments so the original content can be reconstructed after selected fragment loss.

    7

    Does replication replace backup?

    No. Replication maintains current copies, while backups provide historical recovery after events such as deletion or corruption.

    8

    What is RPO?

    Recovery Point Objective defines the acceptable amount of data loss measured in time.

    9

    What is RTO?

    Recovery Time Objective defines the target time for restoring an acceptable level of service after disruption.

    10

    Why are checksums important for durability?

    Checksums help detect stored or transferred data that no longer matches the expected content.

    11

    Why must backups be tested?

    A completed backup does not prove that the required data, credentials, encryption keys, and application can be restored successfully.

    12

    What comes after durability?

    The next topic is checksums, followed by metadata, uploads and multipart transfer, inverted indexes, and full-text search.

    Key Takeaway

    Durability ensures that acknowledged data remains intact and recoverable under a defined failure model. Build durability through appropriate write acknowledgment, persistent logging, replication or erasure coding, checksums, integrity verification, repair, snapshots, versioning, and protected backups. Separate durability from availability and consistency, distribute redundant data across meaningful failure domains, define RPO and RTO, protect encryption keys, and test actual restoration. Replication keeps additional current copies, but historical backups and verified recovery procedures remain necessary for accidental deletion, corruption, faulty updates, and disasters.