Block, file and object storage
Block, File, and Object Storage
Learn how block, file, and object storage organize and expose data, how their access models differ, and how to select storage based on latency, sharing, scale, metadata, update patterns, consistency, durability, security, and operational requirements.
Introduction
Storage is a foundational system-design decision. Applications store databases, operating-system volumes, uploaded documents, videos, backups, logs, reports, machine-learning datasets, and many other forms of data.
These workloads do not all require the same access model. A database may require low-level random reads and writes. A shared application may expect directories and normal file operations. A media platform may need to store billions of independently addressable objects.
The three major storage models are:
- Block storage: Exposes raw addressable blocks to a host or operating system.
- File storage: Exposes files and directories through a hierarchical file system.
- Object storage: Stores complete objects identified by keys and accessed through an API.
Core idea: Block storage exposes blocks, file storage exposes files and directories, and object storage exposes objects through keys and APIs. The correct choice depends on how the application must access, update, share, search, protect, and scale its data.
In your System Design curriculum, Block, File, and Object Storage is Topic 6.1 and begins the Storage, Files, Objects & Search Basics module. It is followed by durability, checksums, metadata, uploads and multipart transfer, inverted indexes, and full-text search.
Prerequisites
| # | Prerequisite | Why It Is Needed |
|---|---|---|
| 1 | Basic computer storage | Storage systems ultimately persist bytes on physical or virtual devices. |
| 2 | Files and directories | File storage exposes a hierarchical namespace through file-system operations. |
| 3 | Networking | Remote block, file, and object systems can be accessed across a network. |
| 4 | HTTP and APIs | Object storage is commonly accessed through service APIs. |
| 5 | Databases | Database engines commonly rely on block devices or managed storage abstractions. |
| 6 | Security fundamentals | Stored data requires authentication, authorization, encryption, retention, and audit controls. |
Understanding the Storage Stack
An application can interact with storage through several abstraction layers.
Application
|
+-- Database API
|
+-- File-system API
|
+-- Object-storage API
|
v
Operating system or storage client
|
v
Local or network storage service
|
v
Physical storage devices
The selected abstraction determines which operations the application can perform and which responsibilities belong to the application, operating system, or storage service.
Block Storage
Block storage divides a storage volume into fixed-size addressable blocks. The storage layer reads or writes blocks without inherently understanding files, directories, database rows, images, or document formats.
Block device
+---------+---------+---------+---------+
| Block 0 | Block 1 | Block 2 | Block 3 |
+---------+---------+---------+---------+
| Block 4 | Block 5 | Block 6 | Block 7 |
+---------+---------+---------+---------+
An operating system can place a file system on the block device, or a database can manage its own data structures over the available volume.
Block-storage Characteristics
- Data is addressed as blocks
- The block layer does not inherently provide filenames or directories
- A host can format the volume with a file system
- Applications can perform random reads and writes
- Only selected blocks need to be changed during an in-place update
- The attached host commonly controls the file system placed on the volume
Common Block-storage Use Cases
- Database data and transaction-log volumes
- Virtual-machine boot and data disks
- Container persistent volumes
- File-system volumes
- Low-latency transactional workloads
- Applications requiring frequent random updates
File Storage
File storage organizes data into files and directories. Applications access content through paths and file-level operations.
/
+-- courses/
| +-- system-design/
| +-- chapter-01.pdf
| +-- chapter-02.pdf
|
+-- learners/
| +-- learner-1042/
| +-- certificate.pdf
|
+-- media/
+-- lesson-01.mp4
The file system manages filenames, directory structures, metadata, and file-level access operations.
File-storage Characteristics
- Hierarchical files and directories
- Path-based access
- File metadata such as size and timestamps
- File and directory permissions
- Shared access through supported network file protocols
- Partial reads and writes through file offsets
- Familiar file-system semantics for applications
Common File-storage Use Cases
- Shared application directories
- User home directories
- Document-management systems
- Content-production workflows
- Legacy applications that require file paths
- Shared configuration or application assets
- Workloads requiring file locking or partial file updates
Object Storage
Object storage stores content as independently addressable objects. Each object generally contains data, associated metadata, and a key or identifier.
Bucket or container
|
+-- Key:
| courses/system-design/lesson-01.mp4
|
+-- Object data:
| video bytes
|
+-- Metadata:
content type
content length
checksum
creation time
custom attributes
Object storage is normally accessed through an API rather than through ordinary local file-system operations.
Object-storage Characteristics
- Objects are addressed through keys
- Objects can contain rich system and custom metadata
- Storage is commonly organized into buckets or containers
- Namespaces commonly appear hierarchical while remaining key-based
- Objects are accessed through APIs
- Complete-object replacement is common for updates
- The model is suitable for large quantities of unstructured data
Common Object-storage Use Cases
- User-uploaded images and videos
- Course documents and learning resources
- Backups and archives
- Application logs
- Data-lake datasets
- Static website assets
- Generated reports and exports
- Machine-learning training data
Block vs File vs Object Storage
| Area | Block Storage | File Storage | Object Storage |
|---|---|---|---|
| Storage unit | Block | File | Object |
| Namespace | Block addresses | Hierarchical directories and paths | Bucket or container with object keys |
| Common access | Device or volume interface | File-system interface | Service API |
| Application view | Raw volume | Files and folders | Objects and metadata |
| Partial update | Block-level changes | File-offset changes | Complete-object replacement is common |
| Metadata model | Limited at the raw block layer | File-system metadata | Object and custom metadata |
| Typical workload | Databases and virtual disks | Shared files and directory-based applications | Media, backups, uploads, logs, and archives |
Similar Content, Different Interfaces
The same logical content can be stored using different models, but the application interface and operational behaviour change.
Video content as block storage:
Stored inside a file system
on an attached volume.
Video content as file storage:
Stored at a shared path such as
/media/course-42/lesson-7.mp4
Video content as object storage:
Stored under an object key such as
courses/42/lessons/7/video.mp4
The correct model depends on how the content must be written, shared, replaced, distributed, protected, retained, and retrieved.
Update Semantics
Block Update
Existing volume data:
[Block A] [Block B] [Block C]
Modified region:
Update Block B
Result:
[Block A] [New Block B] [Block C]
File Update
Open file
Move to byte offset
Write changed bytes
Close file
Object Update
Read or prepare replacement object
Upload new object content
Replace object associated with key
Update rule: Object storage should not be treated as a normal shared disk without verifying the gateway or compatibility layer. Object APIs and file systems can provide substantially different update, locking, rename, and metadata semantics.
Namespace Design
File Namespace
/tenant-17/courses/course-42/lesson-7/video.mp4
Each path component is part of a directory hierarchy.
Object Namespace
tenant-17/courses/course-42/lesson-7/video.mp4
The separators can make the key appear hierarchical, but the storage service can treat the complete value as an object key.
Good Object-key Qualities
- Stable over the object's lifecycle
- Unique within the required scope
- Independent of mutable display names
- Safe to log under the applicable policy
- Compatible with expected prefix-based organization
- Not dependent on untrusted user input alone
Object Metadata
Object metadata describes the stored content.
{
"objectKey": "courses/42/lessons/7/video.mp4",
"contentType": "video/mp4",
"contentLength": 52428800,
"checksum": "stored-checksum-value",
"courseId": "42",
"lessonId": "7",
"visibility": "private"
}
Metadata can support:
- Content-type handling
- Download naming
- Integrity verification
- Lifecycle rules
- Retention classification
- Search indexing
- Access-control decisions
- Operational diagnostics
Metadata limits, mutability, search behaviour, and supported fields depend on the object-storage service.
Store File Metadata in a Database
Applications commonly store large file bytes in object storage while keeping business metadata in a relational database.
CREATE TABLE course_assets
(
asset_id BIGINT PRIMARY KEY,
tenant_id BIGINT NOT NULL,
course_id BIGINT NOT NULL,
lesson_id BIGINT NULL,
object_key VARCHAR(500) NOT NULL,
original_file_name VARCHAR(255) NOT NULL,
content_type VARCHAR(100) NOT NULL,
content_length BIGINT NOT NULL,
checksum_value VARCHAR(200) NOT NULL,
asset_status VARCHAR(30) NOT NULL,
created_at TIMESTAMP NOT NULL,
CONSTRAINT uq_course_assets_object_key
UNIQUE
(
tenant_id,
object_key
),
CONSTRAINT ck_course_asset_length
CHECK (content_length >= 0)
);
The database row represents the application's business record. The object key identifies the stored binary content.
Reference rule: Keep a stable storage identifier in the database rather than storing temporary download links as permanent object identities.
Direct and Shared Access
| Requirement | Commonly Suitable Direction |
|---|---|
| A database requires a low-level data volume | Block storage |
| Several servers require a shared directory tree | File storage |
| A web application stores uploaded media through APIs | Object storage |
| An application requires normal file locking and partial writes | File storage |
| A virtual machine requires a boot disk | Block storage |
| A platform stores backups and long-lived exports | Object storage |
This table provides common architectural directions, not universal rules. Validate the actual storage product's behaviour and workload requirements.
Object Upload Flow
Client selects file
|
v
Application validates upload request
|
v
Application creates upload authorization
|
v
Client uploads object
|
v
Storage verifies or stores content
|
v
Application verifies completion
|
v
Metadata record becomes active
The architecture can route upload bytes through the application or authorize a client to upload directly to object storage.
Object Download Flow
Client requests asset
|
v
Application authenticates caller
|
v
Application authorizes object access
|
v
Application streams content
or creates temporary access
|
v
Client downloads object
Possession of an object key should not automatically grant access to private content.
HTTP Object Upload Example
PUT /course-assets/courses/42/lesson-7/video.mp4 HTTP/1.1
Host: storage.example.com
Content-Type: video/mp4
Content-Length: 52428800
X-Content-Checksum: stored-checksum-value
[BINARY CONTENT]
This is a conceptual example. Authentication, API paths, checksum fields, conditional requests, size limits, and response semantics depend on the selected object-storage service.
PHP Upload-validation Example
<?php
declare(strict_types=1);
function validateUploadedAsset(
array $file,
int $maximumBytes,
array $allowedMimeTypes
): void {
$error =
$file['error'] ??
UPLOAD_ERR_NO_FILE;
if ($error !== UPLOAD_ERR_OK) {
throw new RuntimeException(
'The upload did not complete successfully.'
);
}
$size =
(int)($file['size'] ?? 0);
if ($size <= 0 ||
$size > $maximumBytes) {
throw new RuntimeException(
'The uploaded file size is invalid.'
);
}
$temporaryPath =
(string)($file['tmp_name'] ?? '');
$mimeDetector =
new finfo(
FILEINFO_MIME_TYPE
);
$detectedMimeType =
$mimeDetector->file(
$temporaryPath
);
if (!is_string($detectedMimeType) ||
!in_array(
$detectedMimeType,
$allowedMimeTypes,
true
)) {
throw new RuntimeException(
'The uploaded file type is not allowed.'
);
}
}
File extensions and client-provided content types are insufficient by themselves. Production systems also require authorization, secure filenames, malware controls where appropriate, quotas, rate limits, integrity verification, and protected storage access.
Secure Upload Design
Upload controls can include:
- Authentication
- Object-level authorization
- Maximum file size
- Allowed content types
- Filename normalization
- Generated storage keys
- Integrity checks
- Malware scanning where required
- Tenant quotas
- Rate limiting
- Retention and deletion rules
- Audit records
Object key is copied directly
from an untrusted uploaded filename.
Tenant ID
+
Generated asset ID
+
Approved file extension or content category
Access Control
Storage authorization should reflect the data's business ownership and classification.
Request:
Download learner certificate
Authorization checks:
- Caller is authenticated
- Caller belongs to trusted tenant scope
- Asset record exists
- Caller owns the certificate
or has an approved administrative permission
- Asset status permits download
Storage buckets, shared file systems, and block volumes should follow least privilege. Public access should be enabled only when explicitly required.
Temporary Object Access
Private object access can be granted through a short-lived signed request or by streaming content through an authorized application endpoint.
Application verifies permission
|
v
Application creates temporary access
|
v
Client receives bounded access
|
v
Temporary authorization expires
The duration, allowed operation, object key, content headers, and other conditions should be as restricted as practical.
Encryption
Storage security can protect data:
- In transit between clients and storage
- At rest on storage infrastructure
- Through managed or customer-controlled encryption keys
- Through application-level encryption where required
Encryption does not replace authorization, key management, auditability, backups, retention, or secure deletion.
Replication Is Not Backup
Several copies can improve availability and tolerance of infrastructure failures, but repeated copies can also reproduce accidental deletion or corruption.
| Capability | Primary Purpose |
|---|---|
| Replication | Maintain additional copies for availability and fault tolerance |
| Snapshot | Capture storage state at a point in time |
| Versioning | Retain previous object generations according to policy |
| Backup | Provide an independently recoverable copy |
| Archive | Retain data for long periods using an appropriate access tier |
Object Versioning
Object key:
courses/42/lesson-7/notes.pdf
Version 1:
Original notes
Version 2:
Corrected notes
Version 3:
Latest published notes
Versioning can help recover previous object content, but it increases stored data and requires retention, lifecycle, and deletion rules.
Lifecycle Management
Storage lifecycle policies can move or remove data according to age, state, type, or business retention rules.
New object
|
v
Frequently accessed tier
|
v
Infrequently accessed tier
|
v
Archive tier
|
v
Delete after approved retention period
Lifecycle rules should reflect legal, audit, recovery, and business requirements. Lower-cost tiers can have different access time and retrieval charges.
Storage Cost Dimensions
Storage cost can include:
- Provisioned capacity
- Consumed storage
- Read and write operations
- Data retrieval
- Network transfer
- Snapshots and retained versions
- Replication
- Performance units
- Backup and archive retention
Compare the complete access and lifecycle pattern rather than storage capacity alone.
Performance Dimensions
| Dimension | Meaning |
|---|---|
| Latency | Time required to complete one storage operation |
| Throughput | Amount of data transferred during a period |
| IOPS | Number of input and output operations completed during a period |
| Concurrency | Number of simultaneous storage operations |
| Object or file size | Amount of data handled by one item or operation |
| Access pattern | Sequential, random, whole-object, partial, or metadata-heavy access |
Large-file Uploads
Very large objects can require multipart or resumable transfer.
Large file
|
v
Split into numbered parts
|
v
Upload parts independently
|
v
Retry failed parts
|
v
Request completion
|
v
Storage assembles final object
Multipart transfer is covered in detail later in this module.
Integrity Verification
Applications can calculate and store a checksum to verify that retrieved or transferred data matches the expected content.
sha256sum course-video.mp4
Checksum algorithms, server-side validation, multipart checksums, and metadata fields depend on the selected storage service. Checksums are covered in detail later in this module.
Consistency Expectations
Storage systems can define different visibility and update semantics for reads, writes, overwrites, listings, metadata, and replicas.
The application should verify:
- When a newly written object becomes readable
- When an overwrite becomes visible
- Whether listing reflects the latest changes
- How concurrent writes are resolved
- Whether conditional writes are supported
- How cached content is invalidated
Do not infer one storage service's consistency semantics from another service's behaviour.
Conditional Object Updates
Where supported, a conditional write can prevent one client from overwriting content that changed after it was read.
Client reads object version V1
Another client writes version V2
First client attempts update:
Write only if current version is still V1
Result:
Conditional write fails
instead of overwriting V2 silently
This is the object-storage equivalent of optimistic concurrency for a supported object operation.
Orphaned Objects
Storage and database updates normally do not participate in one ordinary local transaction.
Upload object succeeds
|
v
Database insert fails
|
v
Object exists without application record
The reverse can also occur:
Database row commits
|
v
Object upload fails
|
v
Application record points to missing content
Possible workflow controls include:
- Pending metadata state
- Upload completion verification
- Idempotent finalization
- Cleanup of abandoned uploads
- Reconciliation jobs
- Deletion queues
Safe Upload-finalization Workflow
1. Create pending asset record.
2. Generate storage key.
3. Upload object.
4. Verify object exists,
size matches,
and integrity checks pass.
5. Mark asset active.
6. Clean up incomplete uploads
after a documented expiry period.
The exact flow depends on whether uploads pass through the application or go directly from the client to storage.
Deletion Workflow
Deletion requested
|
v
Authorize deletion
|
v
Mark business record pending deletion
|
v
Delete or queue storage deletion
|
v
Verify outcome
|
v
Finalize metadata state
Deletion design should account for retention holds, object versions, replicas, backups, caches, audit requirements, and retryable failures.
Course-platform Storage Example
| Platform Data | Candidate Storage | Reason |
|---|---|---|
| Relational course database volume | Block storage | Supports the database engine's low-level storage operations |
| Shared legacy document directory | File storage | Applications require shared hierarchical file paths |
| Lesson videos and documents | Object storage | Content is accessed as independently identifiable uploaded objects |
| Generated learner certificates | Object storage | Certificates can be stored as immutable or versioned objects |
| Temporary report-processing workspace | Block or file storage | The processing tool can require local or file-system operations |
| Backups and exports | Object storage | Content can follow retention and lifecycle policies |
Content-delivery Architecture
Learner
|
v
Application authorization
|
v
Content-delivery layer
|
v
Object storage
Frequently downloaded static content can be distributed through a content delivery layer while private access remains governed by the application's authorization contract.
Caching rules must account for updates, deletion, privacy, tenant scope, and time-limited authorization.
Search Is a Separate Concern
Storage systems preserve content, but they do not automatically provide the application-level full-text search experience required by users.
Stored document
|
v
Extract text and metadata
|
v
Build search index
|
v
Execute user query
|
v
Return authorized search results
Inverted indexes and full-text search are covered later in this module.
Storage Observability
Useful storage metrics include:
- Stored capacity
- Object or file count
- Request rate
- Read and write latency
- Throughput
- Error and timeout rate
- Upload failure rate
- Integrity-verification failures
- Abandoned multipart uploads
- Orphaned metadata or objects
- Lifecycle transition count
- Deletion backlog
- Data-transfer volume
- Storage and retrieval cost
Storage Troubleshooting Workflow
- Identify the storage model and service.
- Identify the affected file, volume, bucket, or object key safely.
- Confirm authentication and authorization.
- Confirm network connectivity and endpoint configuration.
- Check capacity, quota, and object-size limits.
- Check read, write, and list error responses.
- Check storage latency and throughput.
- Verify metadata and content type.
- Verify checksum or integrity evidence.
- Check object version and lifecycle state.
- Check application metadata for stale references.
- Check caches and content-delivery behaviour.
- Compare runtime behaviour with the documented storage contract.
Common Storage-design Mistakes
Choosing Storage Only by Capacity
Access model, latency, throughput, updates, sharing, consistency, security, lifecycle, and cost are also important.
Treating Object Storage as a Normal File System
Object APIs can differ from file systems in partial updates, renames, locking, directories, and concurrent access.
Storing Large Uploads Directly in a Relational Row by Default
Large binary content can increase database size, backup volume, replication work, and query overhead. Evaluate an object-storage reference model.
Using Mutable Display Names as Object Keys
Renaming a course, lesson, or user can require expensive storage-key changes and broken references.
Trusting Uploaded Filenames
Untrusted filenames can create collisions, path problems, misleading file types, or unsafe output behaviour.
Making Private Content Public
Public storage access can bypass application authorization and expose tenant or user content.
Using Permanent Download Links for Sensitive Content
Long-lived links can be copied and reused outside the intended access context.
Assuming Replication Is a Backup
Replication can reproduce accidental deletion or corruption across copies.
Ignoring Orphaned Objects
Failed workflows can leave stored objects without metadata or metadata pointing to missing objects.
Ignoring Stored Versions and Multipart Parts
Old versions and abandoned upload parts can continue consuming storage and increasing cost.
Ignoring Retrieval and Network Cost
A low storage price can be offset by frequent operations, retrieval, and data-transfer charges.
Assuming Every Storage System Has the Same Consistency
Read, overwrite, listing, metadata, and replication semantics must be verified for the selected service.
Recommended Test Cases
| Test | Expected Evidence |
|---|---|
| Block-volume restart | Required persisted data remains available after approved restart testing |
| Shared file access | Authorized application instances can access the same file hierarchy |
| Object upload | The object, metadata, content type, and size are stored correctly |
| Maximum upload size | Oversized content is rejected safely |
| Unsupported content type | The upload is rejected before publication |
| Unauthorized download | Private content is not disclosed |
| Checksum verification | Corrupted or incomplete content is detected |
| Concurrent overwrite | Conditional-update or version policy prevents silent lost updates |
| Database failure after upload | The orphan is detected and cleaned according to policy |
| Upload failure after metadata creation | The pending record is retried, expired, or cleaned safely |
| Lifecycle transition | The object moves or expires according to the approved rule |
| Restore test | Required data can be recovered from the approved backup process |
Storage Best Practices
Recommended Practices
- Select storage according to the required access model.
- Use block storage for suitable database, virtual-disk, and random-update workloads.
- Use file storage when applications require shared paths and file-system semantics.
- Use object storage for suitable uploads, media, backups, logs, datasets, and archives.
- Store business metadata separately from large binary content where appropriate.
- Use stable generated object keys.
- Do not trust client filenames or content types.
- Authenticate and authorize every private upload and download.
- Use short-lived, narrowly scoped temporary access.
- Encrypt protected data in transit and at rest.
- Verify uploaded size, type, and integrity.
- Use multipart or resumable transfer for appropriate large objects.
- Define versioning, retention, lifecycle, and deletion rules.
- Maintain recovery procedures in addition to replication.
- Reconcile application metadata with stored content.
- Clean abandoned uploads and orphaned objects.
- Verify consistency and conditional-write semantics.
- Measure latency, throughput, request rate, and data-transfer volume.
- Consider operation, retrieval, replication, and network costs.
- Test failure, recovery, access control, and lifecycle behaviour.
Practice Exercise
Design the storage architecture for your online learning platform.
Requirements
- Store the relational course and quiz database.
- Store lesson videos, PDFs, source-code files, and presentations.
- Store learner profile images.
- Generate and store course certificates.
- Support private course content.
- Support large resumable video uploads.
- Store file metadata in the relational database.
- Generate stable object keys.
- Validate content type and file size.
- Verify uploaded checksums.
- Define object versioning and lifecycle rules.
- Define backup and restore procedures.
- Detect orphaned database rows and objects.
- Define short-lived download authorization.
- Measure storage, retrieval, and network costs.
Storage-selection Template
| Workload | Access Pattern | Candidate Storage | Reason |
|---|---|---|---|
| Relational database | Frequent random reads and writes | Block or managed database storage | Supports database-engine storage operations |
| Lesson videos | Upload once and stream or download many times | Object storage | Provides object-based access and metadata |
| Legacy shared documents | Several application servers need shared paths | File storage | Provides hierarchical shared-file semantics |
| Certificates | Generate and retrieve independently | Object storage | Each certificate can use a stable protected key |
| Backups | Long-term retained recovery content | Object storage | Supports lifecycle and archival strategies |
Frequently Asked Questions
What is block storage?
Block storage exposes fixed-size addressable blocks that a host, file system, or database can organize and update.
What is file storage?
File storage organizes data into files and directories accessed through paths and file-system operations.
What is object storage?
Object storage stores independently addressable objects containing data, metadata, and a unique key or identifier.
Which storage is commonly used for databases?
Database engines commonly use block-oriented or managed database storage suitable for frequent low-level reads and writes.
Which storage is suitable for shared directories?
File storage is suitable when applications require shared files, directories, paths, and file-system semantics.
Which storage is suitable for videos and uploaded documents?
Object storage is commonly suitable for independently addressable media, documents, backups, and other unstructured content.
Can object storage replace a normal file system?
Not automatically. Object storage can have different semantics for paths, partial updates, renames, locks, listing, and concurrent writes.
Should large files be stored in a relational database?
The decision depends on atomicity, backup, access, security, and size requirements. Large application files are commonly stored in object storage with metadata retained in the database.
Is an object key the same as a directory path?
Not necessarily. Separators can create a path-like presentation, while the storage service can treat the complete value as one object key.
Does replication replace backups?
No. Replication and backups address different failure and recovery requirements.
How should private object downloads be protected?
Authenticate and authorize the request, then stream the object through the application or issue narrowly scoped temporary access.
What comes after block, file, and object storage?
The next topic is durability, followed by checksums, metadata, uploads and multipart transfer, inverted indexes, and full-text search.
Key Takeaway
Block storage exposes addressable blocks and is suitable for workloads such as database volumes and virtual disks. File storage exposes shared files and directories through familiar file-system semantics. Object storage exposes objects through keys and APIs and is suitable for large-scale uploads, media, backups, logs, datasets, and archives. Choose storage according to access and update patterns, sharing requirements, latency, throughput, metadata, consistency, security, lifecycle, recovery, and complete operating cost. Keep stable business metadata in the application database where appropriate, protect private content, verify integrity, and reconcile storage operations that cannot share one local transaction with database updates.