durability
Durability
Learn how storage systems preserve acknowledged data across hardware failures, process crashes, corruption, infrastructure outages, accidental deletion, and disasters through replication, erasure coding, checksums, transaction logs, snapshots, backups, versioning, and tested recovery procedures.
Introduction
Storing data successfully is not enough. A reliable storage system must preserve that data after the client receives confirmation that the write has completed.
Durability describes the system's ability to keep acknowledged data intact and recoverable according to its documented failure model.
Durability concerns include:
- Process crashes
- Operating-system failures
- Storage-device failures
- Power interruptions
- Network partitions
- Data corruption
- Availability-zone failures
- Regional disasters
- Accidental deletion
- Malicious modification
- Software defects
Core idea: Durability is meaningful only when the system defines which failures it protects against, when a write is considered successful, how many copies or recovery structures exist, and how the data can be verified and restored.
In your System Design curriculum, Durability is Topic 6.2 under Storage, Files, Objects & Search Basics. It follows block, file, and object storage and precedes checksums, metadata, uploads and multipart transfer, inverted indexes, and full-text search.
Prerequisites
| # | Prerequisite | Why It Is Needed |
|---|---|---|
| 1 | Block, file, and object storage | Durability mechanisms apply differently to volumes, files, and objects. |
| 2 | Basic failure concepts | Durability guarantees are defined against specific failure scenarios. |
| 3 | Replication | Multiple copies can protect against the loss of one storage component. |
| 4 | Transactions | Committed database changes require durable logging and recovery. |
| 5 | Checksums | Integrity verification helps detect corrupted stored or transferred data. |
| 6 | Backup and recovery | Historical recovery protects against failures that replication alone cannot address. |
What Is Durability?
Durability is the guarantee that, after a storage system acknowledges a successful write, the stored data remains protected and recoverable under the failures covered by that system's durability contract.
Client sends write
|
v
Storage system receives data
|
v
Required persistence or redundancy completes
|
v
System acknowledges success
|
v
Covered failure occurs
|
v
Acknowledged data remains present
or can be recovered
Durability is not an unlimited promise. A single-device system, replicated storage service, versioned object store, and geographically distributed backup system protect against different failure scopes.
Durability vs Availability
Durability and availability answer different questions.
| Property | Primary Question | Example |
|---|---|---|
| Durability | Has the acknowledged data remained intact? | The file still exists after a storage-device failure |
| Availability | Can the data be accessed when requested? | The file can be downloaded immediately |
Durable but temporarily unavailable:
The data remains safely stored,
but a network outage prevents access.
Available but insufficiently durable:
The data is currently readable,
but only one vulnerable copy exists.
Design rule: High availability does not automatically prove high durability, and durable data can still become temporarily unavailable.
Durability vs Consistency
| Property | Concern |
|---|---|
| Durability | Whether acknowledged data survives covered failures |
| Consistency | Which valid value or state readers observe |
| Integrity | Whether stored content remains correct and uncorrupted |
A replica can durably retain an older value while another replica stores a newer value. Durability alone does not define which version a reader will observe.
Define the Failure Model
A durability guarantee should identify the failures it is designed to tolerate.
| Failure | Example Protection |
|---|---|
| Application process crash | Write through a durable storage service before acknowledgment |
| Operating-system crash | Persistent journal or transaction log |
| Single drive failure | Replication or erasure coding across devices |
| Server failure | Copies on independent servers |
| Availability-zone failure | Copies or encoded fragments across zones |
| Regional disaster | Geographically separate replica or backup |
| Accidental deletion | Versioning, snapshots, or independent backup |
| Silent corruption | Checksums, verification, and repair from redundant data |
| Malicious encryption or deletion | Protected immutable or isolated recovery copies |
Write Acknowledgment
The moment at which the system confirms a write is central to its durability guarantee.
Weak acknowledgment:
1. Receive write.
2. Keep data only in volatile memory.
3. Return success.
4. Process loses power.
5. Acknowledged data can be lost.
Stronger acknowledgment:
1. Receive write.
2. Persist required recovery information.
3. Create required redundant representation.
4. Return success.
Stronger acknowledgment can require additional storage operations or network round trips, increasing write latency.
Volatile vs Persistent Storage
| Storage | General Property |
|---|---|
| Volatile memory | Content can be lost when power or the process is lost |
| Persistent storage | Designed to retain content across process or power interruption |
| Replicated persistent storage | Maintains redundant data representations across failure domains |
| Backup storage | Maintains recoverable historical copies separately from active data |
Caches improve performance but should not be treated as the only durable copy unless the cache is explicitly designed and configured as a durable storage system.
Write-ahead Logging
A database can use write-ahead logging so that recovery information is made durable before modified data pages are written to their final locations.
Transaction modifies data
|
v
Create transaction-log records
|
v
Persist required log records
|
v
Acknowledge commit
|
v
Write modified data pages later
After a crash, the database can use its transaction log to restore committed operations and remove incomplete work according to its recovery algorithm.
Conceptual Database Transaction
BEGIN;
INSERT INTO enrollments
(
enrollment_id,
learner_id,
course_id,
enrollment_status
)
VALUES
(
:enrollment_id,
:learner_id,
:course_id,
'active'
);
COMMIT;
The database durability configuration determines what persistence work must complete before the commit is acknowledged.
Replication
Data replication maintains multiple copies of stored data, called replicas, across selected storage components or locations.
Client write
|
v
Primary storage
|
+-- Replica A
|
+-- Replica B
|
+-- Replica C
Replication can help protect against drive, server, rack, zone, or regional failure depending on where replicas are placed.
Synchronous Replication
With synchronous replication, the write acknowledgment waits for the required replica confirmations.
Client sends write
|
v
Primary stores write
|
v
Required replica stores write
|
v
Replica acknowledges
|
v
Primary returns success
Potential Benefit
The acknowledged write is already represented on the required participating replicas.
Consideration
Write latency and availability can depend on the slowest required replica or network path.
Asynchronous Replication
With asynchronous replication, the primary can acknowledge a write before a remote replica receives it.
Client sends write
|
v
Primary stores write
|
v
Primary returns success
|
v
Replication continues asynchronously
If the primary is lost before replication completes, recently acknowledged data can be absent from the remote replica, depending on the service guarantee.
Replication Comparison
| Area | Synchronous | Asynchronous |
|---|---|---|
| Acknowledgment | Waits for required replica persistence | Can occur before remote replication completes |
| Write latency | Includes required replication path | Can avoid waiting for remote copy |
| Recent-data exposure | Lower for covered replica failures | Replication lag can expose recent writes |
| Failure handling | Required replica availability can affect writes | Primary can continue while replica catches up |
Quorum-based Persistence
Distributed storage can acknowledge a write after a required subset of replicas confirms it.
Replication factor:
N
Required write acknowledgments:
W
Required read responses:
R
A simplified condition often discussed for overlapping read and write quorums is:
\[ R + W > N \]
This equation alone does not fully describe the consistency or durability of a real system. Replica failures, conflict resolution, repair, timing, and implementation details also matter.
Erasure Coding
Erasure coding divides data into fragments and generates additional parity fragments. The original content can be reconstructed when enough fragments remain available.
Original data
|
v
Split into data fragments
|
v
Generate parity fragments
|
v
Distribute fragments across devices
After selected fragment loss:
Remaining fragments
|
v
Reconstruct original content
Erasure coding can provide storage-efficient redundancy but requires encoding, distribution, reconstruction, and repair work.
Replication vs Erasure Coding
| Area | Replication | Erasure Coding |
|---|---|---|
| Representation | Complete copies of data | Data and parity fragments |
| Recovery | Read from another complete replica | Reconstruct from available fragments |
| Storage overhead | Depends on number of full copies | Depends on data and parity configuration |
| Repair work | Copy from a healthy replica | Read and reconstruct required fragments |
| Typical decision factors | Latency, simplicity, and workload | Capacity efficiency and large-scale durability |
Checksums and Integrity
Durability must protect not only against missing data but also against undetected corruption.
Write data
|
v
Calculate checksum
|
v
Store data and checksum
|
v
Read data later
|
v
Recalculate checksum
|
+-- Match:
| content passed verification
|
+-- Mismatch:
corruption detected
When redundant data exists, the system can use a valid copy or fragment set to repair corrupted content.
File-checksum Example
sha256sum course-video.mp4
Checksums and integrity verification are covered in detail in the next topic after durability.
Scrubbing
Data scrubbing is the periodic process of reading stored data, verifying its integrity, and repairing detected corruption when valid redundant data is available.
Periodic verification job
|
v
Read stored fragment or object
|
v
Verify checksum
|
+-- Valid:
| continue
|
+-- Invalid:
locate healthy copy
repair corrupted data
record event
Without periodic verification, silent corruption can remain undiscovered until recovery is attempted.
Snapshots
A snapshot captures storage state at a specific point in time.
Live volume
|
+-- Snapshot at Time A
|
+-- Snapshot at Time B
|
+-- Snapshot at Time C
Snapshots can support:
- Recovery from accidental modification
- Volume cloning
- Testing against a captured state
- Backup workflows
- Point-in-time recovery strategies
A snapshot stored in the same failure domain can share risks with the source storage. Its independence and recovery value must be verified.
Object Versioning
Object versioning retains previous generations when an object is overwritten or deleted according to the configured policy.
Object key:
courses/42/lesson-7/notes.pdf
Version 1:
Original content
Version 2:
Corrected content
Version 3:
Latest content
Versioning can help recover from accidental overwrite or deletion, but it increases storage consumption and requires lifecycle and retention rules.
Backups
A backup is a timestamped copy of data that can be used to restore lost or damaged information.
Production storage
|
v
Backup process
|
v
Timestamped recovery copy
|
v
Protected backup location
Backup design should define:
- What data is included
- How often copies are created
- Where copies are stored
- How long copies are retained
- How copies are protected
- How restoration is performed
- How restoration is tested
Replication vs Backup
| Area | Replication | Backup |
|---|---|---|
| Primary purpose | Maintain additional current copies | Maintain recoverable historical copies |
| Accidental deletion | Deletion can propagate to replicas | An earlier backup can retain the deleted data |
| Corruption | Corruption can replicate if not detected | A clean earlier copy can support recovery |
| Availability | Can support failover to another copy | Usually requires a restore operation |
| Historical recovery | Not necessarily provided | A primary purpose of backup retention |
Recovery rule: Replication improves redundancy, but it does not replace protected backups. Accidental deletion, corruption, or a faulty application update can affect multiple replicas.
Backup Isolation and Immutability
A recovery copy should be protected from the same identities, errors, or failures that can damage production data.
Protection can include:
- Separate security boundaries
- Restricted deletion permissions
- Retention locks where required
- Immutable backup copies
- Independent credentials
- Geographically separate storage
- Audited recovery operations
Recovery Point Objective
Recovery Point Objective, commonly abbreviated as RPO, describes the acceptable amount of data loss measured in time.
Last recoverable point:
10:00
Failure:
10:15
Potential lost interval:
15 minutes
A smaller RPO requires more frequent or more continuous protection.
Recovery Time Objective
Recovery Time Objective, commonly abbreviated as RTO, describes the target time for restoring an acceptable level of service after a disruption.
Failure detected
|
v
Recovery begins
|
v
Data restored or failover completed
|
v
Service becomes usable
Durability protects data, while RTO addresses how quickly the service must become operational again.
RPO vs RTO
| Objective | Question |
|---|---|
| RPO | How much recent data loss is acceptable? |
| RTO | How quickly must the service or data be restored? |
Failure Domains
Redundant copies should be placed across failure domains appropriate to the required durability target.
Device
|
v
Server
|
v
Rack
|
v
Availability zone
|
v
Region
|
v
Independent backup location
Two copies on one physical device do not protect against that device's failure. Several copies in one building do not protect against the complete loss of that building.
Regional and Geographic Redundancy
Geographic redundancy stores or replicates data in locations separated from the primary storage location.
Design questions include:
- Is replication synchronous or asynchronous?
- What recent-data exposure remains?
- Can applications read from the secondary location?
- How is failover initiated?
- How is failback performed?
- Are encryption keys available in the recovery location?
- Are regulatory location requirements satisfied?
Database Durability
Database durability normally combines transaction logs, checkpoints, persistent data pages, recovery processing, and the underlying storage platform.
Application transaction
|
v
Database transaction log
|
v
Durable storage acknowledgment
|
v
COMMIT returned
|
v
Checkpoint and page persistence
|
v
Recovery after failure
Database durability settings can trade write latency for a weaker or stronger acknowledgment guarantee. Review platform-specific configuration before changing the default.
File-storage Durability
File-storage durability can combine protected underlying devices, file-system journaling, redundant servers, snapshots, checksums, and backups.
Application writes file
|
v
File-system operation
|
v
File metadata and data persisted
|
v
Storage redundancy applied
|
v
Backup or snapshot retained
A successful application write does not automatically prove that all buffers have reached durable media. File APIs, flush operations, mount settings, and storage guarantees must be understood.
Object-storage Durability
An object-storage system can protect data by storing redundant complete copies or encoded fragments across several devices or locations.
Object upload
|
v
Validate request and data
|
v
Store redundant representation
|
v
Persist object metadata
|
v
Acknowledge successful upload
|
v
Periodically verify integrity
Object versioning, retention, lifecycle policies, and independent backup can protect against additional risks such as accidental overwrite or deletion.
Database and Object-storage Workflow
An object upload and relational database update normally cannot share one ordinary local transaction.
Upload object succeeds
|
v
Database metadata insert fails
|
v
Durable object exists
without active business record
A resilient workflow can use:
- Pending metadata status
- Generated object identifiers
- Upload verification
- Idempotent finalization
- Abandoned-object cleanup
- Metadata-to-storage reconciliation
Durable Upload Workflow
1. Create pending asset metadata.
2. Generate a stable object key.
3. Upload the object.
4. Verify object existence,
expected size,
and checksum.
5. Mark the asset active.
6. Retain appropriate versions or backups.
7. Reconcile incomplete workflows.
Accidental-deletion Recovery
Object deleted accidentally
|
v
Check object versions
|
+--> Previous version exists:
| restore selected version
|
+--> No retained version:
locate backup
restore object
verify checksum
verify metadata reference
A deletion process should account for versions, snapshots, replicas, backups, caches, search indexes, legal holds, and business metadata.
Restore Testing
A backup is useful only when the required data can be restored correctly.
A restore test should verify:
- The backup can be located
- Decryption keys are available
- The backup can be read
- The required recovery point exists
- The restored data passes integrity checks
- Application references remain valid
- The restored service can start
- Access control remains correct
- Recovery completes within the required objective
Operational rule: Do not treat a successful backup job as proof of recoverability. Perform controlled restore tests and verify the restored data.
Durability Metrics
A durability percentage communicates a statistical design target for data loss over a period. It is not a promise that an individual object can never be lost.
A simplified annual loss expectation can be expressed as:
\[ ExpectedLostObjects = StoredObjects \times AnnualLossProbability \]
Use provider-published definitions when interpreting reported durability. Confirm the measured unit, period, exclusions, service scope, and failure assumptions.
Durability and Data Retention
Durable storage preserves data, while retention rules determine how long the data should remain.
| Concern | Question |
|---|---|
| Durability | Can the stored data survive covered failures? |
| Retention | How long should the stored data or versions be retained? |
| Deletion | When and how should data be removed? |
| Legal hold | Must normal expiration or deletion be suspended? |
| Archival | Can older data move to a lower-access storage tier? |
Keeping data forever is not automatically safer. Unnecessary retention can increase privacy, compliance, operational, and cost risks.
Encryption-key Durability
Encrypted durable content becomes unusable if the required decryption keys are permanently lost.
Durable encrypted backup
|
v
Decryption key unavailable
|
v
Data cannot be restored
Key-management planning should address:
- Key backup or protected redundancy
- Access control
- Rotation
- Recovery-location availability
- Auditability
- Revocation and destruction
Threats to Durable Data
Durability planning should consider:
- Hardware failure
- Software defects
- Operator mistakes
- Expired credentials
- Lost encryption keys
- Ransomware
- Unauthorized deletion
- Silent corruption
- Incorrect lifecycle policies
- Replication lag
- Incomplete backups
- Untested restoration
Durability Observability
Useful durability metrics include:
- Write acknowledgment failures
- Replication lag
- Under-replicated object or block count
- Checksum verification failures
- Repair backlog
- Storage-device failures
- Snapshot success and failure count
- Backup success and failure count
- Backup age
- Restore-test success rate
- Restore duration
- Unrecoverable-object count
- Expired-version count
- Recovery-point age
Alerting
Alert conditions can include:
- Required replica count is not satisfied
- Replication lag exceeds the approved threshold
- Backup did not complete
- No recent valid recovery point exists
- Checksum mismatches increase
- Repair backlog grows
- Restore test fails
- Backup or archive storage approaches capacity
- Retention or immutability policy changes unexpectedly
Durability-review Workflow
- Identify the data and its business importance.
- Define the acknowledged-write semantics.
- Define the failures the system must tolerate.
- Identify storage and replica failure domains.
- Define required replication or erasure coding.
- Define checksum and integrity verification.
- Define snapshot and versioning policies.
- Define backup frequency and retention.
- Define RPO and RTO.
- Protect recovery copies and encryption keys.
- Implement repair and reconciliation.
- Test failure and restoration procedures.
- Monitor replication, backups, integrity, and recovery readiness.
Common Durability Mistakes
Acknowledging before Required Persistence
A failure can remove data that the client was told had been stored successfully.
Keeping Only One Copy
One device, server, or location can remain a single point of data loss.
Placing Every Replica in One Failure Domain
Several copies cannot protect against a complete failure of their shared location or dependency.
Assuming Replication Is Backup
Deletion, corruption, and faulty updates can propagate to replicas.
Never Verifying Checksums
Silent corruption can remain undetected until the content is needed.
Creating Backups without Restore Tests
A completed backup job does not prove that the data, keys, and application can be recovered.
Storing Backups with Production Permissions
The same compromised identity can delete or corrupt both production data and its recovery copies.
Ignoring Replication Lag
A secondary copy can be missing recent writes when failover occurs.
Ignoring Encryption Keys
Durable encrypted data is not recoverable when its required keys are lost.
Keeping Every Version Forever
Unlimited retention increases cost and can violate privacy or deletion requirements.
Confusing Durability with Availability
Data can remain intact while being temporarily inaccessible.
Using Provider Durability Numbers without Understanding Scope
A published target must be interpreted using its documented period, workload, exclusions, failure model, and service configuration.
Recommended Test Cases
| Test | Expected Evidence |
|---|---|
| Process crash before acknowledgment | The client does not receive a false successful result |
| Process crash after acknowledged commit | The committed data remains recoverable |
| Single-device failure | Data is reconstructed or read from healthy redundancy |
| Replica failure | The under-replicated state is detected and repaired |
| Checksum mismatch | Corruption is detected and not returned silently |
| Accidental object overwrite | The required previous version can be recovered |
| Accidental deletion | The object or file can be restored according to policy |
| Backup failure | Monitoring identifies the missing recovery point |
| Restore test | Data is restored and passes integrity validation |
| Lost encryption-key simulation | The key-recovery procedure is validated safely |
| Regional recovery exercise | Documented data and service recovery objectives are measured |
| Metadata-object reconciliation | Missing and orphaned objects are detected |
Durability Best Practices
Recommended Practices
- Define the durability failure model explicitly.
- Acknowledge writes only after required persistence completes.
- Use redundancy appropriate to the required failure domains.
- Choose synchronous and asynchronous replication deliberately.
- Monitor replication lag and under-replicated data.
- Use checksums to detect corruption.
- Periodically verify stored data and repair damaged copies.
- Use snapshots or versions where point-in-time recovery is required.
- Maintain protected historical backups.
- Separate backup security from production access where practical.
- Define RPO and RTO from business requirements.
- Protect encryption keys required for recovery.
- Apply approved retention and deletion policies.
- Reconcile database metadata with files and objects.
- Clean orphaned and abandoned storage safely.
- Test process, device, zone, and recovery failures where approved.
- Perform recurring restore tests.
- Verify restored data using checksums and application checks.
- Monitor backup age, integrity failures, repair backlog, and restore readiness.
- Document ownership for recovery decisions and procedures.
Practice Exercise
Design a durability and recovery plan for your online learning platform.
Requirements
- Classify course videos, documents, certificates, database records, and audit logs.
- Define the acknowledged-write point for each data type.
- Identify device, server, zone, and regional failure requirements.
- Define replication or erasure-coding requirements.
- Store and verify checksums for uploaded content.
- Define object-versioning rules.
- Define snapshot and backup schedules.
- Define recovery retention periods.
- Set RPO and RTO for critical platform data.
- Protect backup credentials and encryption keys.
- Detect orphaned objects and metadata records.
- Test accidental object deletion.
- Test database restoration.
- Test recovery of one learner certificate.
- Record recovery results and corrective actions.
Durability-design Template
| Data | Failure Model | Protection | Recovery Objective |
|---|---|---|---|
| Course database | Process, device, zone, and operator failure | Transaction log, replication, backup, and restore testing | Record approved RPO and RTO |
| Lesson videos | Device corruption, deletion, and regional disruption | Object redundancy, checksum, versioning, and backup | Record approved recovery requirements |
| Learner certificates | Deletion, corruption, and metadata mismatch | Object protection, checksum, and regeneration or backup | Record approved recovery requirements |
| Audit records | Unauthorized deletion and storage failure | Controlled retention, protected copies, and integrity checks | Record approved retention and recovery rules |
| Search index | Index loss or corruption | Rebuild from authoritative source where supported | Record rebuild target |
Frequently Asked Questions
What is storage durability?
Storage durability is the guarantee that acknowledged data remains intact and recoverable under the failures defined by the storage system.
Is durability the same as availability?
No. Durability concerns whether data remains intact, while availability concerns whether the data can be accessed when requested.
How does replication improve durability?
Replication maintains multiple copies so data can survive the loss of a covered storage component or location.
What is synchronous replication?
Synchronous replication waits for required replica persistence before acknowledging the write.
What is asynchronous replication?
Asynchronous replication allows the primary to acknowledge a write before remote replication is complete.
What is erasure coding?
Erasure coding stores data and parity fragments so the original content can be reconstructed after selected fragment loss.
Does replication replace backup?
No. Replication maintains current copies, while backups provide historical recovery after events such as deletion or corruption.
What is RPO?
Recovery Point Objective defines the acceptable amount of data loss measured in time.
What is RTO?
Recovery Time Objective defines the target time for restoring an acceptable level of service after disruption.
Why are checksums important for durability?
Checksums help detect stored or transferred data that no longer matches the expected content.
Why must backups be tested?
A completed backup does not prove that the required data, credentials, encryption keys, and application can be restored successfully.
What comes after durability?
The next topic is checksums, followed by metadata, uploads and multipart transfer, inverted indexes, and full-text search.
Key Takeaway
Durability ensures that acknowledged data remains intact and recoverable under a defined failure model. Build durability through appropriate write acknowledgment, persistent logging, replication or erasure coding, checksums, integrity verification, repair, snapshots, versioning, and protected backups. Separate durability from availability and consistency, distribute redundant data across meaningful failure domains, define RPO and RTO, protect encryption keys, and test actual restoration. Replication keeps additional current copies, but historical backups and verified recovery procedures remain necessary for accidental deletion, corruption, faulty updates, and disasters.